Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

110

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

110 results for “schema”

Learn how ShareScore rates datasets ↗
zenodo40/100

Text-fig. 14. Schema of transversal section of Manilkaroxylon sp. (sample DR2). ap – axial parenchyma, grb – growth ring boundaries, r – rays, v – vessels. in New Fossil Woods From The Paleogene Of Doupovské Hory And České Středohoří Mts. (Bohemian Massif, Czech Republic)

Text-fig. 14. Schema of transversal section of Manilkaroxylon sp. (sample DR2). ap – axial parenchyma, grb – growth ring boundaries, r – rays, v – vessels.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Text-fig. 7. Schema of radial section of T. gypsaceum (sample 98/04). t – tracheid, r – ray, bp – bordered pit, tp – taxodioid pit, cp – cupressoid pit. in New Fossil Woods From The Paleogene Of Doupovské Hory And České Středohoří Mts. (Bohemian Massif, Czech Republic)

Text-fig. 7. Schema of radial section of T. gypsaceum (sample 98/04). t – tracheid, r – ray, bp – bordered pit, tp – taxodioid pit, cp – cupressoid pit.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Text-fig. 10. Schema of transversal section of A. tschemrylica (sample 97/04). v – vessel, r – ray, ar – aggregate ray, grb – growth-ring boundary. in New Fossil Woods From The Paleogene Of Doupovské Hory And České Středohoří Mts. (Bohemian Massif, Czech Republic)

Text-fig. 10. Schema of transversal section of A. tschemrylica (sample 97/04). v – vessel, r – ray, ar – aggregate ray, grb – growth-ring boundary.

opencc-by-4.0Dec 2015View details →
zenodo40/100

Simple dataset obtained as a Wikidata subset from 2022 dump using entity schema about taxon

<p>The subset has been obtained using wdsub version 0.0.33 and the schema:</p> <p>&nbsp;</p> <pre>PREFIX p: &lt;http://www.wikidata.org/prop/&gt; PREFIX ps: &lt;http://www.wikidata.org/prop/statement/&gt; PREFIX prov: &lt;http://www.w3.org/ns/prov#&gt; PREFIX wd: &lt;http://www.wikidata.org/entity/&gt; PREFIX wdt: &lt;http://www.wikidata.org/prop/direct/&gt; start = @&lt;taxon_by_wd_ontology&gt; OR @&lt;taxon_by_identifier&gt; &lt;taxon_by_wd_ontology&gt; { wdt:P31 [wd:Q16521] ; } &lt;taxon_by_identifier&gt; {wdt:P685 . +;} OR # NCBI taxonomy ID {wdt:P846 . +;} OR # GBIF taxon ID {wdt:P3151 . +;} OR # iNaturalist taxon ID {wdt:P3444 . +;} # eBird taxon ID</pre>

opencc-by-4.0Feb 2023View details →
zenodo40/100

Supplemental material of "An annotated whole-genome multilocus sequence typing schema for scalable high resolution typing of Streptococcus pyogenes"

<p>This supplemental material includes the genome assemblies, associated metadata and analysis results for five datasets used to define a publicly available annotated wgMLST schema for <em>S. pyogenes</em> and to evaluate its suitability for high resolution typing. A brief description for each file in the dataset is available in the included README file. Raw sequencing data and sample metadata for the 265 isolates included in Dataset1 have been deposited in the European Nucleotide Archive (ENA) under Project <a href="https://www.ebi.ac.uk/ena/browser/view/PRJEB49967?show=reads">PRJEB49967</a>.</p> <p>The wgMLST schema was created with <a href="https://github.com/B-UMMI/chewBBACA">chewBBACA</a> and is publicly available at <a href="https://chewbbaca.online/species/1/schemas/1">chewie-NS</a>, where a more detailed description of schema creation, annotation and curation can be found.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Stakeholders Access to Data - Standardized Schema

<p>The schema helps to streamline and enhance the understanding of access regimes while facilitating the creation of a machine-readable version for easier automation or embedding into software. By adopting this schema, the authors aim to provide a practical solution for enhancing data access and reuse while improving transparency and accountability in the process.</p>

opencc-by-4.0Sep 2023View details →
zenodo36/100

Genome–scale approach to study the genetic relatedness among Brucella melitensis strains - wgMLST schema for Brucella melitensis

<p><strong>wgMLST schema for <em>Brucella melitensis</em></strong></p> <p>&nbsp;</p> <p><strong>Schema creation</strong></p> <p>The wgMLST schema was created using the&nbsp;60 complete genomes of&nbsp;<em>Brucella melitensis&nbsp;</em>available at&nbsp;<a href="https://enterobase.warwick.ac.uk/species/index/ecoli">NCBI</a>,&nbsp;as of January 2019,&nbsp;with the chewBBACA&nbsp;v2.0.11 suite (<a href="https://github.com/B-UMMI/chewBBACA">https://github.com/B-UMMI/chewBBACA</a>), using a training file generated by Prodigal v2.6.3 from the <em>B. melitensis</em> 16M reference genome (RefSeq Accession NC_003317 and NC_003318).&nbsp;For curation and validation, the wgMLST schema was&nbsp;further populated with 212 additional draft genomes:&nbsp;157 draft genomes (downloaded from NCBI in January 2019) and 55 draft genomes assembled with<a href="https://github.com/B-UMMI/INNUca">&nbsp;INNUca v3.1</a> (PRJEB30030).</p> <p>File &#39;Bmelitensis_wgMLST_2656_schema.tar.gz&#39; contains the&nbsp;wgMLST&nbsp;schema formatted for chewBBACA and includes a total of 2656 loci.</p>

opencc-by-4.0Dec 2019View details →
zenodo36/100

Schema di riferimento geografico amministrativo dei comuni italiani della base di dati MIC su RaDISAN flusso unico, costruito partendo da Comuni_Italiani_ISTAT_30_06_2023 di ISTAT

<p><strong>Comuni_Italiani_ISTAT_MIC&nbsp;</strong>rappresenta l'anagrafica più recente dei comuni italiani è dovrà integrare la georeferenziazione della base di dati MIC su RaDISAN flusso unico. ISTAT fa un aggiornamento semestrale (<a href="https://www.istat.it/it/archivio/6789">https://www.istat.it/it/archivio/6789</a>) delle unità amministrative territoriali.</p>

opencc-by-4.0Oct 2023View details →
zenodo36/100

ir_metadata: An Extensible Metadata Schema for Information Retrieval Experiments

<p>This dataset accompanies our work that introduces a metadata schema for TREC run files based on the PRIMAD model. PRIMAD considers essential components of computational experiments that possibly can affect reproducibility on a conceptual level. We propose to align the metadata annotations to the PRIMAD components. In order to demonstrate the potential of metadata annotations, we curated a dataset with run files derived from experiments with different instantiations of PRIMAD components and annotated these with the corresponding metadata. With this work, we hope to stimulate IR researchers to annotate run files and improve the reuse value of experimental artifacts even further.</p> <p>&nbsp;</p> <p>This archive contains the following data:</p> <ul> <li> <p><strong>demo.tar.xz</strong> : Selected annotated runs files that are used in the Colab demonstration.</p> </li> <li> <p><strong>metadata.zip</strong> : YAML files containing only the metadata annotations for each run.</p> </li> <li> <p><strong>runs.zip</strong> : The entire set of run files with annotations.</p> </li> </ul> <p>&nbsp;</p> <p>The annotated runs result from the following experiments:</p> <ul> <li> <p>Grossman and Cormack @ TREC Common Core 2017 <a href="https://trec.nist.gov/pubs/trec26/papers/MRG_UWaterloo-CC.pdf">Paper</a> |&nbsp;<a href="https://trec.nist.gov/">Source</a></p> </li> <li> <p>Grossman and Cormack @ TREC Common Core 2018 <a href="https://trec.nist.gov/pubs/trec27/papers/MRG_UWaterloo-CC.pdf">Paper</a> | <a href="https://trec.nist.gov/">Source</a></p> </li> <li> <p>Yu et al. @ TREC Common Core 2018 <a href="https://trec.nist.gov/pubs/trec27/papers/h2oloo-CC.pdf">Paper</a> | <a href="https://github.com/castorini/Anserini/blob/master/docs/runbook-trec2018-h2oloo.md">Source</a></p> </li> <li> <p>Yu et al. @ ECIR 2019 <a href="https://link.springer.com/chapter/10.1007/978-3-030-15712-8_26">Paper</a> | <a href="https://github.com/castorini/anserini/blob/master/docs/runbook-ecir2019-ccrf.md">Source</a></p> </li> <li> <p>Breuer et al. @ SIGIR 2020 <a href="https://dl.acm.org/doi/10.1145/3397271.3401036">Paper</a> | <a href="https://zenodo.org/record/3856042">Source</a></p> </li> <li> <p>Breuer et al. @ CLEF 2021 <a href="https://link.springer.com/chapter/10.1007/978-3-030-85251-1_5">Paper</a> | <a href="https://zenodo.org/record/4105885">Source</a></p> </li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo36/100

FAIR Metadata Concepts in DataCite Metadata Schema

<p>Documentation Concepts that support the FAIR Principles are mapped to the DataCite Metadata Schema using json Paths.</p>

opencc-by-4.0Jun 2024View details →
zenodo36/100

FAIRmat Tutorial 9: Plugins: Python schemas and parsers

<p>NOMAD is a research data management platform for materials science. NOMAD Oasis allows you to operate the popular NOMAD service for your own lab, with your rules, and on your resources. You can adopt NOMAD Oasis to implement your institutes data policies and to work with your specific data types and workflows.</p> <p>This tutorial aims to introduce participants to the new plugin mechanism in NOMAD and teach them how to develop and integrate their own Python schemas and parsers to a NOMAD Oasis. Plugins enable you to alter how NOMAD processes data and therefore allow for more powerful customisations than the custom schemas presented in past tutorials. Participants will learn how to enable the conversion of new materials science data formats into NOMAD's standardised and machine-readable format. NOMAD plugins can be contributed to the community to further promote reproducibility and transparency in materials science.</p> <p><strong>Disclaimer:</strong> NOMAD is being continuously developed based on input and feedback from the scientific community. Hence the features, services or interface may have changed since the time of recording of this video. For up-to-date information please consult our latest tutorials and the NOMAD documentation <a href="https://nomad-lab.eu/prod/v1/docs/">https://nomad-lab.eu/prod/v1/docs/</a></p>

opencc-by-4.0May 2023View details →
zenodo36/100

FAIRmat Tutorial 14: Developing schemas and parsers for FAIR computational data storage using NOMAD-Simulations

<p><a href="https://nomad-lab.eu"><u>NOMAD</u></a> is an open-source, community-driven data infrastructure, focusing on materials science data. Originally built as a repository for data from DFT calculations, the NOMAD software can automatically extract data from the output of a large variety of simulation codes. Our previous computation-focused tutorials (<a href="https://fairmat-nfdi.github.io/AreaC-Tutorial-CECAM-2023/"><u>CECAM workshop</u></a>, <a href="https://fairmat-nfdi.github.io/AreaC-Tutorial10_2023/"><u>Tutorial 10</u></a>, and <a href="https://www.fairmat-nfdi.eu/events/fairmat-tutorial-7/tutorial-7-materials"><u>Tutorial 7</u></a>) have highlighted the extension of NOMAD&rsquo;s functionalities to support advanced many-body calculations, classical molecular dynamics simulations, and complex simulation workflows.&nbsp;<br>But how can you utilize this infrastructure and associated suite of tools if your simulation code or method is not yet supported?&nbsp;<strong>This tutorial will provide foundational knowledge for customizing NOMAD to fit the specific needs of your computational research project</strong>. The following provides an outline of the major topics that will be covered:</p> <ul> <li>Introduction to the NOMAD software and repository</li> <li>Working with the NOMAD-Simulations schema plugin</li> <li>Extending NOMAD-Simulations to support custom methods and outputs</li> <li>Creating parser plugins from scratch</li> <li>Extra: Interfacing complex simulation and analysis workflows with NOMAD</li> </ul> <p><strong>Disclaimer:</strong> NOMAD is being continuously developed based on input and feedback from the scientific community. Hence the features, services or interface may have changed since the time of recording of this video. For up-to-date information please consult our latest tutorials and the NOMAD documentation <a href="https://nomad-lab.eu/prod/v1/docs/">https://nomad-lab.eu/prod/v1/docs/</a></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

INNUENDO whole genome and core genome MLST schemas and datasets for Campylobacter jejuni

<p><strong>Dataset</strong></p> <p>Raw reads deposited in the European Nucleotide Archive (ENA) or in the NCBI Sequence Read Archive (SRA) as <em>C. jejuni</em> were retrieved in April 2017. In total 5,691 genomes passed the INNUca v3.1 pipeline have been selected. Additionally, 566 raw reads previously published in <a href="https://www.ncbi.nlm.nih.gov/pubmed/27041390">Kovanen et al., 2016</a>,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/28348829">Llarena et al., 2016</a>,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/pubmed/25232158">Kovanen et al., 2014</a>, <a href="https://www.ncbi.nlm.nih.gov/pubmed/24655229">Kovanen et al., 2014</a> and <a href="http://www.sciencedirect.com/science/article/pii/S0740002016310449?via%3Dihub">Gacia-Sanchez et a., 2017</a> were included. The database also includes 269 <em>C. jejuni</em> belonging to the INNUENDO Sequence Dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>).&nbsp;Genomes were assembled using&nbsp;<a href="https://github.com/INNUENDOCON/INNUca">INNUca v3.1 pipeline</a>&nbsp;and passed the QC.&nbsp;</p> <p>File &#39;Metadata/Cjejuni_metadata.txt&#39; contains metadata information for each strain including country and year of isolation, source classification and taxa of the host, classical pubMLST 7 genes ST and CC classification.&nbsp;</p> <p>The directory &#39;Genomes&#39; contains all the 6,526 INNUca V3.1 assemblies of the strains listed in &#39;Metadata/Cjejuni_metadata.txt&#39;.</p> <p><strong>Schema creation and validation</strong></p> <p>Draft genome assemblies were annotated using Prokka and initial pangenome was defined using Roary. The&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA CreateSchema.py</em></a>&nbsp;was used for creating a whole genome schema starting from roary pangenome. The schema was initially composed by 5,447 loci and has been populated with the&nbsp;6,526 <em>C. jejuni</em>&nbsp;genomes. The quality of the loci has been assessed using&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a>. Loci with single alleles and those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) have been removed. The wgMLST schema has been further curated, excluding all those loci detected as &ldquo;Repeated Loci&rdquo; and loci annotated as &ldquo;non-informative paralogous hit (NIPH/ NIPHEM)&rdquo; or &ldquo;Allele Larger/ Smaller than length mode (ALM/ ASM)&rdquo; by the&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a>&nbsp;engine in more than 1% of the&nbsp;<em>C. jejuni</em>&nbsp;genomes dataset.</p> <p>File &#39;Schema/Cjejuni_wgMLST_2795_schema.tar.gz&#39; contains the&nbsp;wgMLST&nbsp;schema formatted for chewBBACA and includes a total of&nbsp;2,795 loci.</p> <p>File &#39;Schema/Cjejuni_cgMLST_678_listGenes.txt&#39; contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of&nbsp;678 loci and has been&nbsp;defined as the loci present in at least the&nbsp;99.9% of the 6,526 <em>C. jejuni</em>&nbsp;genomes. Genomes have no more than 2% of missing loci.</p> <p>File &#39;Allele_Profles/Cjejuni_wgMLST_alleleProfiles.tsv&#39; contains the wgMLST allelic profile of the 6,526 <em>C. jejuni</em>&nbsp;genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File &#39;Allele_Profles/Cjejuni_cgMLST_alleleProfiles.tsv&#39; contains the cgMLST allelic profile of the 6,526 <em>C. jejuni</em>&nbsp;genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citations</strong></p> <p>The schema are prepared to be used with&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carri&ccedil;o J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166&nbsp;<a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote>

opencc-by-4.0Jul 2018View details →
zenodo36/100

INNUENDO whole genome and core genome MLST schemas and datasets for Escherichia coli

<p><strong>Dataset</strong></p> <p>As reference dataset,&nbsp;2,218&nbsp;public draft or complete genome assemblies and available metadata of&nbsp;<em>Escherichia coli</em>&nbsp;have been downloaded from&nbsp;<a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a>&nbsp;in April 2017. Genomes have been selected on the basis of the ribosomal ST (rST) classification available in&nbsp;<a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a>: from the same rST, genomes have been randomly selected and downloaded. The number of samples for each rST in the final dataset is proportional to those available in&nbsp;<a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a>&nbsp;in April 2017. The dataset includes also&nbsp;119<em> </em>Shiga toxin-producing <em>E.coli</em> genomes assembled with<a href="https://github.com/B-UMMI/INNUca"> INNUca v3.1 </a>belonging to the INNUENDO Sequence Dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>).</p> <p>File &#39;Metadata/Ecoli_metadata.txt&#39; contains metadata information for each strain including source classification, taxa of the hosts, country and year of isolation, serotype, pathotype, classical pubMLST 7 genes ST classification, assembly source/method and Enterobase barcode.&nbsp;</p> <p>The directory &#39;Genomes&#39; contains the 119 INNUca v3.1 assemblies of the strains listed in &#39;Metadata/Ecoli_metadata.txt&#39;. Enterobase assemblies can be downloaded from http://enterobase.warwick.ac.uk/species/ecoli/search_strains using &#39;barcode&#39;.</p> <p><strong>Schema creation and validation</strong></p> <p>The wgMLST schema from&nbsp;<a href="https://enterobase.warwick.ac.uk/species/ecoli/download_data">EnteroBase</a>&nbsp;have been downloaded and curated using&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA AutoAlleleCDSCuration</em></a>&nbsp;for removing all alleles that are not coding sequences (CDS). The quality of the remain loci have been assessed using&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a>&nbsp;and loci with single alleles, those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) and those present in less than 0.5% of the&nbsp;<em>Escherichia</em>&nbsp;genomes in&nbsp;<a href="https://enterobase.warwick.ac.uk/species/index/ecoli">EnteroBase</a>&nbsp;at the date of the analysis (April 2017) have been removed. The wgMLST schema have been further curated, excluding all those loci detected as &ldquo;Repeated Loci&rdquo; and loci annotated as &ldquo;non-informative paralogous hit (NIPH/ NIPHEM)&rdquo; or &ldquo;Allele Larger/ Smaller than length mode (ALM/ ASM)&rdquo; by the&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a>&nbsp;engine in more than 1% of a dataset composed by&nbsp;2,337&nbsp;<em>Escherichia coli</em> genomes.</p> <p>File &#39;Schema/Ecoli_wgMLST_7601_schema.tar.gz&#39; contains the&nbsp;wgMLST&nbsp;schema formatted for chewBBACA and includes a total of 7,601 loci.</p> <p>File &#39;Schema/Ecoli_cgMLST_2360_listGenes.txt&#39; contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of 2,360 loci and has been&nbsp;defined as the loci present in at least the&nbsp;99% of the 2,337&nbsp;<em>Escherichia coli</em> genomes. Genomes have no more than 2% of missing loci.</p> <p>File &#39;Allele_Profles/Ecoli_wgMLST_alleleProfiles.tsv&#39; contains the wgMLST allelic profile of the 2,337&nbsp;<em>Escherichia coli</em> genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File &#39;Allele_Profles/Ecoli_cgMLST_alleleProfiles.tsv&#39; contains the cgMLST allelic profile of the 2,337&nbsp;<em>Escherichia coli</em> genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citations</strong></p> <p>The schema are prepared to be used with&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also:</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carri&ccedil;o J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166&nbsp;<a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote> <p><em>Escherichia coli</em> schema is a derivation of EnteroBase <em>E. coli</em> <a href="http://enterobase.warwick.ac.uk/">EnteroBase</a>&nbsp;wgMLST schema. When using the schema in this repository please cite also:</p> <blockquote> <p>Alikhan N-F, Zhou Z, Sergeant MJ, Achtman M (2018) A genomic overview of the population structure of&nbsp;<em>Salmonella</em>. PLoS Genet 14 (4):e1007261.&nbsp;<a href="https://doi.org/10.1371/journal.pgen.1007261">https://doi.org/10.1371/journal.pgen.1007261</a></p> </blockquote>

opencc-by-4.0Jul 2018View details →
zenodo36/100

INNUENDO whole genome and core genome MLST schemas and datasets for Salmonella enterica

<p><strong>Dataset</strong></p> <p>As reference dataset,&nbsp;4,307 public available draft or complete genome assemblies and available metadata of&nbsp;<em>Salmonella enterica</em>&nbsp;have been downloaded from public repositories (i.e.&nbsp;<a href="https://enterobase.warwick.ac.uk/">EnteroBase</a>,&nbsp;<a href="https://www.ncbi.nlm.nih.gov/">National Center for Biotechnology Information NCBI</a>and&nbsp;<a href="https://www.ebi.ac.uk/">The European Bioinformatics Institute EMBL-EBI</a>; accessed April 2017). The collection includes 1,465&nbsp;<em>S.</em>&nbsp;Enteritidis,&nbsp;2,442 <em>S.</em>Typhimurium, and&nbsp;400 of other frequently isolated serovars in Europe. The dataset includes also 153 <em>S.</em>Typhimurium variant 4,[5],12:i:- collected from different Italian regions between 2012 and 2014 during a surveillance study and&nbsp;129&nbsp;<em>S.</em>&nbsp;Enteritidis belonging to the INNUENDO sequence dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>). The 282 additional genomes were assembled using <a href="https://github.com/B-UMMI/INNUca">INNUca v3.1</a>.</p> <p>File &#39;Metadata/Senterica_metadata.txt&#39; contains metadata information for each strain including source classification, host taxa, year and country of isolation, serotype, classical pubMLST 7 genes ST classification, and source/method of the assembly.&nbsp;</p> <p>The directory &#39;Genomes&#39; contains all the 4,589 assemblies of the strains listed in &#39;Metadata/Senterica_metadata.txt&#39;. Please note that genomes marked as &#39;Enterobase&#39; have been downloaded from Enterobase webpage http://enterobase.warwick.ac.uk.</p> <p><strong>Schema creation and validation</strong></p> <p>The wgMLST schema from&nbsp;<a href="https://enterobase.warwick.ac.uk/species/senterica/download_data">EnteroBase</a>&nbsp;have been downloaded and curated using&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA AutoAlleleCDSCuration</em></a>&nbsp;for removing all alleles that are not coding sequences (CDS). The quality of the remain loci have been assessed using&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a>&nbsp;and loci with single alleles, those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) and those present in less than 0.5% of the&nbsp;<em>Salmonella</em>&nbsp;genomes in&nbsp;<a href="https://enterobase.warwick.ac.uk/species/index/senterica">EnteroBase</a>&nbsp;at the date of the analysis (April 2017) have been removed. The wgMLST schema have been further curated, excluding all those loci detected as &ldquo;Repeated Loci&rdquo; and loci annotated as &ldquo;non-informative paralogous hit (NIPH/ NIPHEM)&rdquo; or &ldquo;Allele Larger/ Smaller than length mode (ALM/ ASM)&rdquo; by the&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a>&nbsp;engine in more than 1% of a dataset composed by&nbsp;4,589 <em>Salmonella</em>&nbsp;genomes.</p> <p>File &#39;Schemas/Senterica_wgMLST_ 8558_schema.tar.gz&#39; contains the&nbsp;wgMLST&nbsp;schema formatted for chewBBACA and includes a total of&nbsp; 8,558 loci.</p> <p>File &#39;Schemas/Senterica_cgMLST_ 3255_listGenes.txt&#39; contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of&nbsp; 3,255 loci and has been&nbsp;defined as the loci present in at least the&nbsp;99% of the 4,589 <em>Salmonella</em>&nbsp;genomes. Genomes have no more than 2% of missing loci.</p> <p>File &#39;Allele_Profles/Senterica_wgMLST_alleleProfiles.tsv&#39; contains the wgMLST allelic profile of the 4,589 <em>Salmonella</em>&nbsp;genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File &#39;Allele_Profles/Senterica_cgMLST_alleleProfiles.tsv&#39; contains the cgMLST allelic profile of the 4,589 <em>Salmonella</em>&nbsp;genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citations</strong></p> <p>The schema are prepared to be used with&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also:</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carri&ccedil;o J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166&nbsp;<a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote> <p><em>Salmonella enterica</em> schema is a derivation of EnteroBase <em>Salmonella </em><a href="http://enterobase.warwick.ac.uk/">EnteroBase</a>&nbsp;wgMLST schema. When using the schema in this repository please cite also:</p> <blockquote> <p>Alikhan N-F, Zhou Z, Sergeant MJ, Achtman M (2018) A genomic overview of the population structure of&nbsp;<em>Salmonella</em>. PLoS Genet 14 (4):e1007261.&nbsp;<a href="https://doi.org/10.1371/journal.pgen.1007261">https://doi.org/10.1371/journal.pgen.1007261</a></p> </blockquote>

opencc-by-4.0Jul 2018View details →
zenodo36/100

INNUENDO whole genome and core genome MLST schemas and datasets for Yersinia enterocolitica

<p><strong>Dataset</strong></p> <p>All the raw reads deposited in the European Nucleotide Archive (ENA) or in the NCBI Sequence Read Archive (SRA) as <em>Y. enterocolitica</em> at the time of the analysis (August 2018) were retrieved using <a href="https://github.com/B-UMMI/getSeqENA">getSeqENA</a>. A total of 252 genomes were successfully assembled using <a href="https://github.com/B-UMMI/INNUca">INNUca v3.1</a>. In addition to public available genomes, the database includes 79 novel <em>Y. enterocolitica</em> strains which belong to the INNUENDO Sequence Dataset (<a href="https://www.ebi.ac.uk/ena/data/view/PRJEB27020">PRJEB27020</a>).&nbsp;</p> <p>File &#39;Metadata/Yenterocolitica_metadata.txt&#39; contains metadata information for each strain including country and year of isolation, source classification, taxon of the host, serotype, biotype, pathotype (according to patho_typing software) and classical pubMLST 7 genes ST according to <a href="http://jcm.asm.org/content/53/1/35.long">Hall et al., 2005</a>.&nbsp;</p> <p>The directory &#39;Genomes&#39; contains all the 331 INNUca V3.1 assemblies of the strains listed in &#39;Metadata/Yenterocolitica_metadata.txt&#39;.</p> <p><strong>Schema creation and validation</strong></p> <p>All the 331 genomes were used for creating the schema using&nbsp;<strong><a href="https://github.com/B-UMMI/chewBBACA">chewBBACA suite</a></strong>. The quality of the loci have been assessed using&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/1.-Schema-Creation"><em>chewBBACA Schema Evaluation</em></a>&nbsp;and loci with single alleles, those with high length variability (i.e. if more than 1 allele is outside the mode +/- 0.05 size) and those present in less than 1% of the genomes have been removed. The wgMLST schema have been further curated, excluding all those loci detected as &ldquo;Repeated Loci&rdquo; and loci annotated as &ldquo;non-informative paralogous hit (NIPH/ NIPHEM)&rdquo; or &ldquo;Allele Larger/ Smaller than length mode (ALM/ ASM)&rdquo; by the&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki/2.-Allele-Calling"><em>chewBBACA Allele Calling</em></a>&nbsp;in more than 1% of a dataset.</p> <p>File &#39;Schema/Yenterocolitica_wgMLST_ 6344_schema.tar.gz&#39; contains the&nbsp;wgMLST&nbsp;schema formatted for chewBBACA and includes a total of&nbsp; 6,344 loci.</p> <p>File &#39;Schema/Yenterocolitica_cgMLST_ 2406_listGenes.txt&#39; contains the list of genes from the wgMLST schema which defines the cgMLST schema. The cgMLST schema consists of&nbsp; 2,406 loci and has been&nbsp;defined as the loci present in at least the&nbsp;99% of the 331<em> Y. enterocolitica</em> genomes. Genomes have no more than 2% of missing loci.</p> <p>File &#39;Allele_Profles/Yenterocolitica_wgMLST_alleleProfiles.tsv&#39; contains the wgMLST allelic profile of the 331&nbsp;<em>Y. enterocolitica</em> genomes of the dataset. Please note that missing loci follow the annotation of chewBBACA Allele Calling software.</p> <p>File &#39;Allele_Profles/Yenterocolitica_cgMLST_alleleProfiles.tsv&#39; contains the cgMLST allelic profile of the 331&nbsp;<em>Y. enterocolitica</em> genomes of the dataset. Please note that missing loci are indicated with a zero.</p> <p><strong>Additional citation</strong></p> <p>The schema are prepared to be used with&nbsp;<a href="https://github.com/B-UMMI/chewBBACA/wiki"><strong>chewBBACA</strong></a>. When using the schema in this repository please cite also:</p> <blockquote> <p>Silva M, Machado M, Silva D, Rossi M, Moran-Gilad J, Santos S, Ramirez M, Carri&ccedil;o J. chewBBACA: A complete suite for gene-by-gene schema creation and strain identification. 15/03/2018. M Gen 4(3): doi:10.1099/mgen.0.000166&nbsp;<a href="http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166">http://mgen.microbiologyresearch.org/content/journal/mgen/10.1099/mgen.0.000166</a></p> </blockquote>

opencc-by-4.0Jul 2018View details →
zenodo36/100

Text-fig. 13. Schema of typical rays in G. ortenburgense (sample 89/04 and 90/04). in New Fossil Woods From The Paleogene Of Doupovské Hory And České Středohoří Mts. (Bohemian Massif, Czech Republic)

Text-fig. 13. Schema of typical rays in G. ortenburgense (sample 89/04 and 90/04).

opencc-by-4.0Dec 2015View details →
zenodo36/100

Text-fig. 15. Schema of observed rays of Manilkaroxylon sp. (sample DR2). in New Fossil Woods From The Paleogene Of Doupovské Hory And České Středohoří Mts. (Bohemian Massif, Czech Republic)

Text-fig. 15. Schema of observed rays of Manilkaroxylon sp. (sample DR2).

opencc-by-4.0Dec 2015View details →
zenodo36/100

UNIC Example implementations of the metadata schema for interpreting corpora

<p>Four example implementations of the metadata schema for interpreting corpora, namely the European Parliament Interpreting Corpus v2.0 (Russo et al. 2012), the interpreted subcorpus of the <em>Dolmetschen im Krankenhaus</em> (&lsquo;Interpreting in Hospitals&rsquo;) corpus v1.1&nbsp;(B&uuml;hrig et al. 2012), the Speech Corpus of Interpreted Premier Press Conferences v1.0 (Liu 2023), and the Belgian Covid Sign language corpus v1.1 (Vandeghinste et al. 2022).&nbsp;</p>

opencc-by-4.0Aug 2024View details →
zenodo36/100

UNIC Example implementations of the metadata schema for interpreting corpora

<div> <p>Four example implementations of the metadata schema for interpreting corpora, namely the European Parliament Interpreting Corpus v2.0 (Russo et al. 2012), the interpreted subcorpus of the&nbsp;<em>Dolmetschen im Krankenhaus</em> (&lsquo;Interpreting in Hospitals&rsquo;) corpus v1.1&nbsp;(B&uuml;hrig et al. 2012), the Speech Corpus of Interpreted Premier Press Conferences v1.0 (Liu 2023), and the Belgian Covid Sign language corpus v1.1 (Vandeghinste et al. 2022).&nbsp;</p> </div>

opencc-by-4.0Aug 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record