Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
599
datasets available to search
ShareScore release 0.7.1
Dataset results
599 results for “metadata”
S71 | CECSCREEN | HBM4EU CECscreen: Screening List for Chemicals of Emerging Concern Plus Metadata and Predicted Phase 1 Metabolites
<p>This is the collection associated with list S71 CECSCREEN HBM4EU CECscreen: Screening List for Chemicals of Emerging Concern Plus Metadata and Predicted Phase 1 Metabolites<strong> </strong>on the NORMAN Suspect List Exchange.</p> <p><a href="https://www.norman-network.com/nds/SLE/">https://www.norman-network.com/nds/SLE/</a></p> <p>CECScreen is part of the HBM4EU project (coord. UBA) > WP16 "emerging chemicals" (lead INRA, JP Antignac/L Debrauwer) > Task 16.1 (lead IRAS, J Vlanderen / R Vermeulen) > Main contributor (J Meijer) > Involved Partners (M Lamoree, T Hamers, S Hutinet, A, Covaci, C Huber, M Krauss, DI Walker, EL Schymanski). Further details in Meijer et al (2021) DOI: <a href="https://doi.org/10.1016/j.envint.2021.106511">10.1016/j.envint.2021.106511</a>. Dataset DOI: <a href="https://doi.org/10.5281/zenodo.3956586">10.5281/zenodo.3956586</a>.</p> <p>Update 23/7/2020 (v0.1.1): updated MetFrag files to remove elements causing errors (Os, Pd, Ag, Be). Update 8 Nov 2022 (v0.1.2) removed new lines in several synonyms as detected at BioHackEU22.</p>
Ethnic and Migrant Minorities (EMM) Survey Registry: All metadata records
<p>The <a href="https://ethmigsurveydatahub.eu/emmregistry/">Ethnic and Migrant Minorities (EMM) Survey Registry</a> is a free online tool that allows users to search for and learn about existing quantitative surveys undertaken with EMM (sub)populations conducted in 34 European countries, from 2000 onwards, through compiled survey-level metadata.</p> <p>The first version was produced by a team led by CEE (Sciences Po, CNRS) and jointly funded through the COST Action 16111 – ETHMIGSURVEYDATA (a network of more than 200 European researchers active in the ethnic and migration studies field), the Horizon 2020 infrastructure project SSHOC (within Task 9.2 on Ethnic and Migration Studies, within Work Package 9 on Data Communities) and the project FAIRETHMIGQUANT (an Open Science project funded by the French Agence Nationale de la Recherche, ANR).</p> <p>This specific record includes the metadata for 2,120 survey records as .dta, .sav and .csv files published on the Registry, as of 31.07.2025.</p>
Dataset and program scripts for the reproducibility of the hierarchical data structure file. Related to the manuscript entitled: Hierarchical Representation of Measurement Data, Metrological Uncertainty and Metadata for Calibrated Battery Tests
<p>We present an interoperable hierarchical data representation for battery tests, leading to improved scalability of data transmission and enhanced data accessibility and comprehensibility for both human interpretation and machine processing. The hierarchical data format includes the raw trace electrical measurement data, the metrological calibration and uncertainty data, the metadata such as experimental settings, instruments and software versions, as well as post-processed data such as electrochemical model fit parameters. This data representation allows repetition of the battery test under the exact same conditions such that identical results are achieved within defined error bounds. This is in line with the general F.A.I.R. data approach and provides repeatability and traceability in the battery value chain. As an application of the hierarchical data representation, we show the classification of cells as pass/fail being performed with quantitative confidence levels. We demonstrate the complete workflow of establishing the hierarchical data structure for electrochemical impedance spectroscopy (EIS), starting from metrological traceability of the calibration and uncertainty analysis towards the storage of the structured data as a single integrated file that preserves the hierarchical data format.</p>
UNIC Templates for uploading corpus metadata v1.11
<p>The UNIC platform (https://unic.dipintra.it) accepts a JSON file for uploading corpus metadata based on the template here. Alternatively, use the spreadsheet template to input the corpus metadata and convert the resulting .xlsm file to JSON using this application at https://huggingface.co/spaces/nannanliu/UNIC_metadata_conversion. When opening the Excel spreadsheet template, please enable Macros, which will automatically validate your input in the columns. Please do not change the order of the columns because they are embedded with code. To add elements and components not included by the UNIC schema, create new columns after the existing ones.</p>
Dominant contribution of Asgard archaea to eukaryogenesis (2024) Tobiasson, V., Koonin, E. PROCESSED DATA AND METADATA
<h1>Main data deposit for "Dominant contribution of Asgard archaea to eukaryogenesis". </h1> <p>Victor Tobiasson, Jacob Luo, Yuri I Wolf, Eugene V Koonin</p> <p>Computational Biology Branch, Division of Intramural Research, National Library of Medicine, National Institutes of Health, Bethesda, MD 20894, USA</p> <p><strong>The Origin of eukaryotes is one of the key problems in evolutionary biology. The demonstration that the Last Eukaryotic Common Ancestor (LECA) already contained the mitochondrion, an endosymbiotic organelle derived from an alphaproteobacterium, and the discovery of Asgard archaea, the closest archaeal relatives of eukaryotes inform and constrain evolutionary scenarios of eukaryogenesis. We undertook a comprehensive analysis of the origins of the core eukaryotic genes tracing to the LECA within a rigorous statistical framework centered around evolutionary hypotheses testing using constrained phylogenetic trees. The results reveal dominant contributions of Asgard archaea to the origin of most of the conserved eukaryotic functional systems and pathways. A limited contribution from Alphaproteobacteria was identified, primarily relating to the energy transformation systems and Fe-S cluster biogenesis, whereas ancestry from other bacterial phyla was scattered across the eukaryotic functional landscape, without consistent trends. These findings suggest a model of eukaryogenesis in which key features of eukaryotic cell organization evolved in the Asgard ancestor, followed by the capture of the Alphaproteobacterial endosymbiont, and augmented by numerous but sporadic horizontal acquisition of genes from other bacteria both before and after endosymbiosis. </strong></p> <div> <div> <div>Version 0.3, updated 180325</div> <div> </div> <div> </div> <div>Main data repository for:</div> <div>Dominant contribution of Asgard archaea to eukaryogenesis (2024) </div> <div>Tobiasson, V., Koonin, E.</div> <div> </div> <div>Contains all final parsed data from the main Eukaryogenesis project </div> <div>investigating the evolutionary ancetries of eukaryotic protein families. </div> <div> </div> <div>Currently (non-static) available at: </div> <div>https://www.biorxiv.org/content/10.1101/2024.10.14.618318v2</div> <div>https://assets-eu.researchsquare.com/files/rs-5352492/v1/2f9c68ae-cf3e-420a-8d29-867b6fb1a878.pdf</div> <div> </div> <div>All code used to generate the data present within this repository available at: </div> <div>https://github.com/VictorTobiasson/eukgen </div> <div> </div> <div> </div> <div>### General information</div> <div> </div> <div>To identify associations between prokaryotic and eukaryotic protein families, separate</div> <div>hidden Markov model (HMM) databases for prokaryotes and eukaryotes were constructed </div> <div>using a custom, cascaded, sequence-to-profile clustering pipeline, implemented using </div> <div>mmseqs2, followed by a multistep data-reduction and multiple sequence alignment (MSA) </div> <div>procedure to generate HMM profiles using hhsuite. </div> <div> </div> <div>A prokaryotic database of 37 million protein sequences was curated from prokaryotic </div> <div>genomes obtained from the NCBI GenBank in November 2023 and supplemented with proteins </div> <div>extracted from 146 Asgard genome assemblies. To avoid inclusion of genes present only </div> <div>within a narrow subset of species, possibly resulting from horizontal transfer from </div> <div>eukaryotes post LECA, we reconstructed the “soft-core” pangenome for each of the 26 </div> <div>curated prokaryotic taxonomic classes. These pangenomes include only those genes that </div> <div>are present in at least 67% of the families within each class of Bacteria and Archaea. </div> <div>The initial eukaryotic database consisted of 30 million protein sequences from 993 </div> <div>species taken from EukprotV3 and cleaned using mmseqs2 to remove likely prokaryotic </div> <div>contaminants. </div> <div> </div> <div>Both databases were clustered and MSAs constructed for all non, singleton clusters </div> <div>and HMM profiles created. The resulting eukaryotic HMM dataset was queried against </div> <div>the prokaryotic dataset using hhblits to identify sets of homologous protein sequences. </div> <div>Each eukaryotic cluster and all its significant prokaryotic hits constituted an individual</div> <div> sequence set, hereinafter referred to as an Eukaryotic/Prokaryotic Orthologous Cluster </div> <div>(EPOC). The EPOCs constitute groups of homologous proteins from eukaryotes and prokaryotes </div> <div>(each EPOC contains a unique set of eukaryotic proteins, but some clusters of prokaryotic </div> <div>proteins can be present in multiple EPOCs) that were used for phylogenetic tree </div> <div>construction, annotation, and evolutionary hypothesis testing. </div> <div> </div> <div>To infer the most likely prokaryotic ancestry of the eukaryotic proteins in each EPOC, </div> <div>rather than relying on the tree topology directly, we employed a probabilistic approach </div> <div>for evolutionary hypothesis testing using constraint trees. We exhaustively sampled all </div> <div>arrangements of likely sister clades and obtained Expected Likelihood Weights (ELW) for </div> <div>the set of possible sister clade models. As the ELW metric is analogous to model selection </div> <div>confidence, here we take it to be proportional to the probability of a sampled prokaryotic </div> <div>clade to be the true sister group of the given eukaryotic clade among a set of competing </div> <div>sister clades. For each EPOC, our analysis dynamically accounts for long branch outliers </div> <div>and is robust to phylogenetically non-homogenous clades. This analysis is further capable </div> <div>of resolving eukaryotic paraphyly, treating each eukaryotic clade within a EPOC as a </div> <div>single datapoint for downstream analysis. Our resulting data contains EPOCs annotated </div> <div>using profiles generated from KEGG Orthology Groups (KOGs), each with an MSA generated </div> <div>using muscle5, a maximum likelihood tree inferred using IQtree2 and associated ELW values </div> <div>for all candidate prokaryotic sister phyla. The analysis of prokaryotic ancestry was </div> <div>performed only for those eukaryotic clades that included more than 5 distinct taxonomic </div> <div>labels, with at least one coming from Amorphea and one from Diaphoretickes, the two </div> <div>expansive eukaryotic clades considered to represent either the first or the second </div> <div>bifurcation in the evolution of eukaryotes. Thus, these clades likely represent genes </div> <div>mapping back to the LECA.</div> <div> </div> <div>For further details please see main publication or contact</div> <div>victor.tobiasson@nih.gov</div> <div>eugene.koonin@nih.gov</div> <div> </div> <div> </div> <div>### Included files</div> <div> </div> <div>Unless otherwise stated all files contained are tab separated and utf-8 encoded </div> <div>with the first row containing header information. </div> <div>All data entries encoding lists are “|” (pipe) separated. </div> <div>Fields without data values are filled with string entries of “none”.</div> <div> </div> <div>--- Databases ---</div> <div>euk72_ep.tar.gz</div> <div>prok2311_as.tar.gz</div> <div>Prok2311As_final_clusters.tsv</div> <div>Euk72Ep_final_clusters.tsv</div> <div>prok2311_as.hmmDB.tar.gz</div> <div>euk72_ep.hmmDB.tar.gz</div> <div> </div> <div>--- Annotation and Curation ---</div> <div>NCBI_taxonomy_species_addendum.tsv</div> <div>NCBI_taxonomy_class_addendum.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>KEGG_category_mapping.tsv</div> <div>KEGG_metadata.tsv</div> <div> </div> <div>--- EPOC data ---</div> <div>EPOC_data.tar.gz</div> <div>EPOC_annotation_KEGG.tsv</div> <div>EPOC_data.tsv</div> <div>EPOC_data.pangenomes_s10.tsv</div> <div>EPOC_data.pangenomes_s25.tsv</div> <div>EPOC_data.pangenomes_s67.tsv</div> <div>EPOC_data.GTDB.tsv</div> <div> </div> <div># euk72_ep.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files </div> <div>constituting the initial eukaryotic mmseqs2 database with taxonomy annotation. </div> <div>Constructed from a pre-selected list of 72 eukaryotic proteomes downloaded from </div> <div>NCBI as well as a “clean” version of Eukprot, lacking highly prokaryotic-like </div> <div>contaminant sequences. </div> <div> </div> <div># prok2311_as.tar.gz</div> <div>Gunzip-ed .tar archive containing a single directory with 10 files constituting the </div> <div>initial prokaryotic mmseqs2 database with taxonomy annotation. Constructed from </div> <div>47545 complete genomes retrieved from NCBI in November 2023. </div> <div> </div> <div># prok2311_as.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted </div> <div>from prok2311_as non--singleton clusters, contains 26286 profiles.</div> <div> </div> <div># euk72_ep.hmmDB.tar.gz</div> <div>Gunzip-ed .tar archive containing 6 files. Comprises an HHSuite Databse formatted </div> <div>from euk72_ep non-singleton clusters, contains 1631704 profiles.</div> <div> </div> <div># NCBI_taxonomy_species_addendum.tsv</div> <div>Taxonomy mapping file with manually curated ‘class’ level annotation for poorly </div> <div>annotated species. </div> <div> </div> <div>taxid: NCBI taxid</div> <div>proposed_class_id: Manually assigned NCBI taxid</div> <div>proposed_class_label: NCBI class name</div> <div>org_name: NCBI organism name</div> <div> </div> <div># NCBI_taxonomy_class_addendum.tsv</div> <div>Class revision file mapping poorly populated class level entries to higher order </div> <div>manually curated labels. Also includes information for small classes with shallow </div> <div>taxonomy which are deleted from the EPOC analysis at the level of tree construction.</div> <div> </div> <div>taxid: NCBI taxid</div> <div>ncbi_class: NCBI taxid of rank corresponding to ‘class’ following manual </div> <div>amendment as per NCBI_taxonomy_species_addendum.tsv</div> <div>revised_class_id: Manually assigned NCBI taxid of rank corresponding to ‘class’</div> <div>revised_class_label: Proposed cleartext name of manually revised revised_class_id </div> <div> </div> <div># Euk72Ep_Prok2311As_final_classes.tsv</div> <div>Final taxonomy at NCBI rank ‘class’ following revisions for all sequences in Euk72Ep or </div> <div>Prok2311As. These taxonomic labels are used for EPOC tree annotation. </div> <div> </div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya, </div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of manually revised NCBI rank ‘class’ identifier for annotation</div> <div> </div> <div># Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Final taxonomy at GTDB rank ‘phylum’ transferred using marker genes from GTDB release 220</div> <div> </div> <div>acc: mmseqs database header in either prok2311_as or euk72_ep databases</div> <div>taxid: NCBI taxid for organism</div> <div>superkingdom: Top level NCBI taxonomy classification Bacteria, Archaea or Eukarya, </div> <div>used to define Eukaryotic outgroups in EPOC analysis</div> <div>class: Cleartext name of assigne GTDB phylum</div> <div> </div> <div># Prok2311As_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the </div> <div>final clusters used for HMM creation </div> <div> </div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div> </div> <div># Euk72Ep_final_clusters.tsv</div> <div>Cluster mapping file for accessions within the initial Prok2311A database to the </div> <div>final clusters used for HMM creation</div> <div> </div> <div>cluster_acc: cluster representative</div> <div>acc: cluster member</div> <div> </div> <div># EPOC_data.tar.gz</div> <div>Gunzip-ed directory containing 16035 EPOC folders. Each folder is named corresponding </div> <div>to the eukaryotic cluster representative which generated its profile as an ID </div> <div>Matches the tree_name field in EPOC_data_prok2311As.tsv</div> <div>contains the following files:</div> <div> </div> <div><EPOC_ID>.merged.fasta: sequences for all members of the EPOC</div> <div><EPOC_ID>.merged.fasta.leaf_mapping: tsv separated file containing taxonomy and tree reduction data</div> <div><EPOC_ID>.merged.fasta.muscle: main cropped MSA for tree generation </div> <div><EPOC_ID>.merged.fasta.muscle.iqtree: IQtree2 output from tree generation</div> <div><EPOC_ID>.merged.fasta.muscle.treefile.annot: annotated newick tree file with final tree</div> <div><EPOC_ID>.merged.tree_data.tsv: final parsed tree data with columns matching EPOC_data_prok2311As.tsv</div> <div> </div> <div>EPOCs with more than one possible eukaryotic sister phyla also contains </div> <div>a folder "constraint_analysis" with constraint tree information used for </div> <div>ELW value calculation. </div> <div> </div> <div># EPOC_data.tsv</div> <div>Main resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) </div> <div>based on pangenomes defined as including 10% of species per class. This is the main</div> <div>data to be used for genereting the core dataset and for data visualistation</div> <div>Contains information regarding tree breakdown, LCA membership and phylogenetic </div> <div>distances between all detected LCAs. Equivalent to the stacked dataframes from all </div> <div>EPOC directories in EPOC_data </div> <div> </div> <div>tree_name: unique index for each EPOC </div> <div>euk_clade_rep: unique index for each annotated eukaryotic clade within each tree_name</div> <div>euk_clade_size: number of original sequences represented by euk_clade_rep</div> <div>euk_clade_weight: metric for taxonomic purity for each euk_clade_rep</div> <div>euk_leaf_clade: boolean indicating whether euk_clade_rep contains a single leaf</div> <div>euk_LCA: lowest taxa spanning all members in euk_clade_rep</div> <div>euk_scope: list of all taxonomic classes in euk_clade_rep</div> <div>euk_scope_len: length of euk_scope list</div> <div>prok_clade_rep: unique index for each annotated prokaryotic clade for each euk_clade_rep</div> <div>prok_clade_size: number of original sequences represented by prok_clade_rep</div> <div>prok_clade_weight: metric for taxonomic purity for each prok_clade_rep</div> <div>prok_leaf_clade: boolean indicating whether prok_clade_rep contains a single leaf</div> <div>prok_taxa: lowest taxa spanning all members in prok_clade_rep</div> <div>dist: tree-distance from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>top_dist: graph-distance (node-distance) from lowest tree node containing all members of prok_clade_rep to lowest tree node containing all members of euk_clade_rep</div> <div>raw_stem_length: tree-distance from lowest tree node containing the union of all members of prok_clade_rep and euk_clade_rep to the tree node containing all members of euk_clade_rep</div> <div>median_euk_leaf_dist: median value for all tree distances from the tree node containing all members of euk_clade_rep to the individual leaves</div> <div>stem_length: raw_stem_length/median_euk_leaf_dist</div> <div>logL: log likelihood of best constraint tree constructed</div> <div>deltaL: log likelihood difference between constraint tree for prok_clade_rep and best constraint tree constructed</div> <div>bp-RELL: validation metric from IQtree -trees, see iqtree.org</div> <div>bp-RELL_accept: as above</div> <div>p-KH: as above</div> <div>p-KH_accept: as above</div> <div>p-SH: as above</div> <div>p-SH_accept: as above</div> <div>c-ELW: as above</div> <div>c-ELW_accept: as above</div> <div>p-AU: as above</div> <div>p-AU_accept: as above</div> <div> </div> <div># EPOC_data.pangenomes_s10.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>based on pangenomes defined as including 10% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.pangenomes_s25.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>based on pangenomes defined as including 25% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.pangenomes_s67.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>based on pangenomes defined as including 67% of species per class.</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.GTDB.tsv</div> <div>Resulting data from all Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>under revised taxonomy from GTDB based on data from Euk72Ep_Prok2311As_final_classes.GTDB.tsv</div> <div>Identical file structure to EPOC_data.tsv</div> <div> </div> <div># EPOC_data.alpha_replicates.tsv</div> <div>Resulting data from 20 repetitions of Eukaryotic/Prokaryotic Orthologous Clusters (EPOCs) calculated </div> <div>from a subset of Alphaproteobacterial-derived EPOCs. </div> <div>Identical file structure to EPOC_data.tsv with the addition of:</div> <div> </div> <div>rep: indicating technical replicate number, 0-19</div> <div> </div> <div># EPOC_annotation_KEGG.tsv</div> <div>Parsed HHblits output of HMM profiles generated from KEGG KOGs (KEGG Orthologous Groups) </div> <div>against eukaryotic profiles constituting each EPOC</div> <div> </div> <div>Query: query name equal to tree_name from EPOC_data</div> <div>Target: target name equal to kogid in KEGG_category_mapping and KEGG_metadata</div> <div>Prob: data from HHblits, see https://github.com/soedinglab/hh-suite/wiki</div> <div>E-value : as above</div> <div>P-value : as above</div> <div>Score: as above</div> <div>SS: as above</div> <div>Cols: as above</div> <div>Identities: as above</div> <div>Similarity: as above</div> <div>Sum_probs: as above</div> <div>Query-HMM-start: as above</div> <div>Query-HMM-end: as above</div> <div>Template-HMM-start: as above</div> <div>Template-HMM-end: as above</div> <div>Template_columns: as above</div> <div>Template_Neff : as above</div> <div>Pairwise_cov: calculated pairwise coverage from Query and Target start and end</div> <div>Description: category_name from KEGG_category_mapping</div> <div> </div> <div># KEGG_category_mapping.tsv</div> <div>Mapping of relevant KOG identifiers to their higher order categories as </div> <div>"Maps" "Modules" or "Reactions" as per KEGG see https://www.kegg.jp/kegg/pathway.html</div> <div> </div> <div>kogid: unique KOG identifier</div> <div>category_id: KEGG map, module, or reaction number</div> <div>category_name: cleartext name for KOG identifier</div> <div> </div> <div># KEGG_metadata.tsv</div> <div>File mapping KOGs to BRITE classification and to additional databases of chemical properties.</div> <div> </div> <div>kogid: unique KOG identifier</div> <div>name: cleartext name for KOG identifier</div> <div>brite_A: list of BRITE-A sets including KOG</div> <div>brite_B: list of BRITE-A sets including KOG</div> <div>brite_C: list of BRITE-A sets including KOG</div> <div>EC: list of Enzyme commission numbers associated with KOG, see https://enzyme.expasy.org/</div> <div>TC: list of transporter classification numbers associated with KOG, see https://www.tcdb.org/</div> <div>RN: list of KEGG reaction numbers associated with KOG</div> <div>CA: list of CAZY numbers associated with KOG, see http://www.cazy.org/</div> <div>GO: list of GO terms associated with KOG, see https://geneontology.org/</div> </div> <div> </div> </div>
Example Microscopy Metadata JSON files produced using Micro-Meta App to document example microscopy experiments performed at individual core facilities
<p>Example <strong>Microscopy Metadata </strong>(Microscope.JSON and Settings.JSON)<strong> files </strong>produced using<strong> <a href="https://wu-bimac.github.io/MicroMetaApp.github.io/">Micro-Meta App</a> </strong>to document the <strong>Hardware Specifications</strong> of example Microscopes and the <strong>Image Acquisition Settings</strong> utilized to acquire example images as listed in the table below.</p> <blockquote> <p>For each facility, the dataset contains two JSON files:</p> <ol> <li><strong>Microscope.JSON file</strong> (e.g., 01_marcello_uliverpool_cci_zeiss_axioobserz1_lsm710.json)</li> <li><strong>Settings.JSON file</strong> (indicated with the name of the image and with the _AS suffix)</li> </ol> </blockquote> <p><strong>Micro-Meta App was</strong> developed as part of a <strong>global community initiative</strong> including the <a href="http://www.4dnucleome.org/"><strong>4D Nucleome (4DN)</strong> </a>Imaging Working Group, <strong>BioImaging North America (BINA)</strong> <a href="https://www.bioimagingna.org/qc-dm-wg">Quality Control and Data Management Working Group</a>, and <strong>QUAlity and REProducibility for Instrument and Images in Light Microscopy</strong> (<a href="https://quarep.org/"><strong>QUAREP-LiMi</strong></a>), to extend the <strong>Open Microscopy Environment (OME)</strong> <a href="https://www.openmicroscopy.org/Schemas/Documentation/Generated/OME-2016-06/ome.html">data model</a>.</p> <blockquote> <p>The works of this <strong>global community effort</strong> resulted in multiple publications featured on a recent <strong>Nature Methods FOCUS ISSUE </strong>dedicated to <a href="https://www.nature.com/collections/djiciihhjh">Reporting and reproducibility in microscopy</a>.</p> </blockquote> <blockquote> <p><strong>Learn More!</strong> For a thorough description of <strong>Micro-Meta App</strong> consult our recent <a href="https://doi.org/10.1038/s41592-021-01315-z">Nature Methods</a> and <a href="https://doi.org/10.1101/2021.05.31.446382">BioRxiv.org</a> publications!</p> </blockquote> <p> </p> <table> <tbody> <tr> <td><strong>Nr.</strong></td> <td><strong>Manufacturer</strong></td> <td><strong>Model</strong></td> <td><strong>Tier</strong></td> <td><strong>Εxperiment Type</strong></td> <td><strong>Facility Name</strong></td> <td><strong>Department and Institution</strong></td> <td><strong>URL</strong></td> <td><strong>References</strong></td> </tr> <tr> <td>1</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer Z1 (with LSM 710 scan head)</strong></td> <td>1</td> <td>3D visualization of superhydrophobic polymer-nanoparticles</td> <td>Centre for Cell Imaging (CCI)</td> <td>University of Liverpool</td> <td>https://cci.liv.ac.uk/equipment_710.html</td> <td>Upton et al., 2020</td> </tr> <tr> <td>2</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer (Axiovert 200M)</strong></td> <td>2</td> <td>Μeasurement of illumination stability on Chinese Hamster Ovary cells expressing Paxillin-EGFP</td> <td>Advanced BioImaging Facility (ABIF).</td> <td>McGill University</td> <td>https://www.mcgill.ca/abif/equipment/axiovert-1</td> <td>Kiepas et al., 2020</td> </tr> <tr> <td>3</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer Z1 (with Spinning Disk)</strong></td> <td>2</td> <td>Immunofluorescence imaging of cryosection of Mouse kidney</td> <td>Imagerie Cellulaire; Quality Control managed by Miacellavie (https://miacellavie.com/)</td> <td>Centre de recherche du Centre Hospitalier Université de Montréal (CR CHUM), University of Montreal</td> <td>https://www.chumontreal.qc.ca/crchum/plateformes-et-services (the web site is for all core facilities, not specifically for the core facility hosting this microscope)</td> <td>Pilliod et al., 2020</td> </tr> <tr> <td>4</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Imager Z2 (with Apotome)</strong></td> <td>2</td> <td>Immunofluorescence imaging of mitotic division in Hela cells using </td> <td>Bioimaging Unit</td> <td>Newcastle University</td> <td>https://www.ncl.ac.uk/bioimaging/</td> <td>Watson et al., 2020</td> </tr> <tr> <td>5</td> <td><strong>Carl Zeiss Microscopy</strong></td> <td><strong>Axio Observer Z1</strong></td> <td>2</td> <td>Fluorescence microscopy of human skin fibroblasts from Glycogen Storage Disease patients.</td> <td>Life Imaging Center (LIC)</td> <td>Centre for Integrative Signalling Analysis (CISA), University of Freiburg</td> <td>https://miap.eu/equipments/sd-i-abl/</td> <td>Hannibal et al., 2020</td> </tr> <tr> <td>6</td> <td><strong>Leica Microsystems</strong></td> <td><strong>DMI6000B</strong></td> <td>2</td> <td>3D immunofluorescence imaging rhinovirus infected macrophages </td> <td>IMAG'IC Confocal Microscopy Facility</td> <td>Institut Cochin, CNRS, INSERM, Université de Paris</td> <td>https://www.institutcochin.fr/core_facilities/confocal-microscopy/cochin-imaging-photonic-microscopy/organigram_team/10054/view</td> <td>Jubrail et al., 2020</td> </tr> <tr> <td>7</td> <td><strong>Leica Microsystems</strong></td> <td><strong>DM5500B</strong></td> <td>2</td> <td>Immunofluorescence analysis of the colocalization of PML bodies with DNA double-strand breaks</td> <td>Bioimaging Unit</td> <td>Edwardson Building on the Campus for Ageing and Vitality, Newcastle University</td> <td>https://www.ncl.ac.uk/bioimaging/equipment/leica-dm5500/#overview</td> <td>da Silva et al., 2019; Nelson et al., 2012<br> </td> </tr> <tr> <td>8</td> <td><strong>Leica Microsystems</strong></td> <td><strong>DMI8-CS (with TCS SP8 STED 3X)</strong></td> <td>2</td> <td>Live-cell imaging of N. benthamiana leaves cells-derived protoplasts</td> <td>Center for Advanced Imaging (CAi)</td> <td>School of Mathematics/Natural Sciences, Heinrich-Heine-Universität Düsseldorf</td> <td>https://www.cai.hhu.de/en/equipment/super-resolution-microscopy/leica-tcs-sp8-sted-3x</td> <td>Singer et al., 2017; Hänsch et al., 2020</td> </tr> <tr> <td>9</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti</strong></td> <td>2</td> <td>Immunofluorescence analysis of the cytoskeleton structure in COS cells</td> <td>Advanced Imaging Center (AIC)</td> <td>Janelia Research Campus, Howard Hughes Medical Institute</td> <td>https://www.janelia.org/support-team/light-microscopy/equipment</td> <td>Abdelfattah et al., 2019; Qian et al., 2019; Grimm et al., 2020</td> </tr> <tr> <td>10</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti-E (HCA)</strong></td> <td>2</td> <td>Τime-lapse analysis of the bursting behavior of amine-functionalized vesicular assemblies</td> <td>Light Microscopy Facility (IALS-LIF)</td> <td>Institute for Applied Life Sciences, University of Massachusetts at Amherst</td> <td>https://www.umass.edu/ials/light-microscopy</td> <td>Fernandez et al., 2020</td> </tr> <tr> <td>11</td> <td><strong>Nikon Instruments/Coleman laboratory (customized)</strong></td> <td><strong>TIRF HILO Epifluorescence light Microscope (THEM)/ Eclipse Ti</strong></td> <td>2</td> <td>Single-particle tracking of Halo-tagged PCNA in Lox cells</td> <td>Coleman laboratory</td> <td>Anatomy and Structural Biology Department, The Albert Einstein College of Medicine</td> <td>https://einsteinmed.org/faculty/12252/robert-coleman/</td> <td>Drosopoulos et al., 2020</td> </tr> <tr> <td>12</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti (with Andor Dragon Fly Spinning Disk)</strong></td> <td>2</td> <td>Investigation of the 3D structure of cerebral organoids</td> <td>Montpellier Resources Imagerie</td> <td>Centre de Recherche de Biologie cellulaire de Montpellier (MRI-CRBM), CNRS, Univerity of Montpellier</td> <td>https://www.mri.cnrs.fr/en/optical-imaging/our-facilities/mri-crbm.html</td> <td>Ayala-Nunez et al., 2019</td> </tr> <tr> <td>13</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti2</strong></td> <td>2</td> <td>Ιmmunofluorescence imaging of cryosections of mouse hearth myocardium </td> <td>Neuroscience Center Microscopy Core</td> <td>Neuroscience Center, University of North Carolina</td> <td>https://www.med.unc.edu/neuroscience/core-facilities/neuro-microscopy/</td> <td>Aghajanian et al., 2021</td> </tr> <tr> <td>14</td> <td><strong>Nikon Instruments</strong></td> <td><strong>Eclipse Ti2</strong></td> <td>2</td> <td>Live-cell imaging of bacterial cells expressing GFP-PopZ</td> <td>Microscopy Resources on the North Quad (MicRoN)</td> <td>Harvard Medical School </td> <td>https://micron.hms.harvard.edu/</td> <td>Lim and Bernhardt 2019; Lim et al., 2019</td> </tr> <tr> <td>15</td> <td><strong>Olympus/Biomedical Imaging Group (customized)</strong></td> <td><strong>TIRF Epifluorescence Structured light Microscope (TESM)/IX71</strong></td> <td>3</td> <td>3D distribution of HIV-1 in the nucleus of human cells</td> <td>Biomedical Imaging Group</td> <td>Program in Molecular Medicine, University of Massachusetts Medical School</td> <td>https://trello.com/b/BQ8zCcQC/tirf-epi-fluorescence-structured-light-microscope</td> <td>Navaroli et al., 2012</td> </tr> <tr> <td>16</td> <td><strong>Olympus/Computer Vision Laboratory (customized)</strong></td> <td><strong>3D BrightField Scanner/IX71</strong></td> <td>3</td> <td>Transmitted light brightfield visualization of swimming spermatocytes</td> <td>Laboratorio Nacional de Microscopia Avanzada (LNMA) and Computer Vision Laboratory of the Institute of Biotechnology</td> <td>Universidad Nacional Autonoma de Mexico (UNAM)</td> <td>https://lnma.unam.mx/wp/</td> <td>Pimentel et al., 2012; Silva-Villalobos et al., 2014</td> </tr> </tbody> </table> <p><strong>Getting started</strong></p> <p>Use these videos to get started with using Micro-Meta App after installation into OMERO and downloading the example data files:</p> <ol> <li><a href="https://vimeo.com/562022222">Video 1</a></li> <li><a href="https://vimeo.com/562022281">Video 2</a></li> </ol> <p><strong>More information</strong></p> <blockquote> <p>For full information on how to use Micro-Meta App please utilize the following resources:</p> <ol> <li>Micro-Meta App <a href="https://wu-bimac.github.io/MicroMetaApp.github.io/">website</a></li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/index.html">Full documentation</a></li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/docs/intro/installation.html">Installation</a> instructions</li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/docs/tutorials/index.html#step-by-step-instructions">Step-by-Step Instructions</a></li> <li><a href="https://micrometaapp-docs.readthedocs.io/en/latest/docs/tutorials/VideoTutorials.html#micro-meta-app-video-tutorials">Tutorial Videos</a></li> </ol> </blockquote> <p><strong>Background</strong></p> <p>If you want to learn more about the importance of <strong>metadata and quality contro</strong>l to ensure full <strong>reproducibility, quality and scientific value</strong> in light microscopy, please take a look at our recent publications describing the development of community-driven light <strong>4DN-BINA-OME Microscopy Metadata</strong> specifications <a href="https://doi.org/10.1038/s41592-021-01327-9">Nature Methods</a> and <a href="https://doi.org/10.1101/2021.04.25.441198">BioRxiv.org</a> and our <a href="https://arxiv.org/abs/1910.11370">overview manuscript</a> entitled <strong>A perspective on Microscopy Metadata: data provenance and quality control</strong>.</p> <p> </p> <p> </p>
Example Microscopy Metadata JSON files produced using Micro-Meta App to document the acquisition of example images using a custom-built TIRF Epifluorescence Structured Illumination Microscope
<p><strong>Example Microscopy Metadata JSON files produced using the <a href="https://wu-bimac.github.io/MicroMetaApp.github.io/">Micro-Meta App</a> documenting an example raw-image file acquired using the custom-built TIRF Epifluorescence Structured Illumination Microscope.</strong></p> <p>For this use case, which is presented in Figure 5 of <a href="http://doi: https://doi.org/10.1101/2021.05.31.446382">Rigano et al., 2021</a>, Micro-Meta App was utilized to document:</p> <p>1) The <strong>Hardware Specifications</strong> of the custom build TIRF Epifluorescence Structured light Microscope (TESM; <a href="https://www.pnas.org/content/109/8/E471.long">Navaroli et al., 2010</a>) developed, built on the basis of the based on Olympus IX71 microscope stand, and owned by the Biomedical Imaging Group (http://big.umassmed.edu/) at the Program in Molecular Medicine of the University of Massachusetts Medical School. Because TESM was custom-built the most appropriate documentation level is <strong>Tier 3</strong> (<em>Manufacturing/Technical Development/Full Documentation</em>) as specified by the <a href="https://doi.org/10.5281/zenodo.4710731">4DN-BINA-OME</a> Microscopy Metadata model (<a href="https://doi.org/10.1101/2021.04.25.441198">Hammer et al., 2021</a>).</p> <p>The TESM Hardware Specifications are stored in: <strong>Rigano et al._Figure 5_UseCase_Biomedical Imaging Group_TESM.JSON</strong></p> <p>2) The <strong>Image Acquisition Settings</strong> that were applied to the TESM microscope for the acquisition of an example image (FSWT-6hVirus-10minFIX-stk_4-EPI.tif.ome.tif) obtained by Nicholas Vecchietti and Caterina Strambio-De-Castillia. For this image, TZM-bl human cells were infected with HIV-1 retroviral three-part vector (FSWT+PAX2+pMD2.G). Six hours post-infection cells were fixed for 10 min with 1% formaldehyde in PBS, and permeabilized. Cells were stained with mouse anti-p24 primary antibody followed by DyLight488-anti-Mouse secondary antibody, to detect HIV-1 viral Capsid. In addition, cells were counterstained using rabbit anti-Lamin B1 primary antibody followed by DyLight649-anti-Rabbit secondary antibody, to visualize the nuclear envelope and with DAPI to visualize the nuclear chromosomal DNA.</p> <p>The Image Acquisition Settings used to acquire the FSWT-6hVirus-10minFIX-stk_4-EPI.tif.ome.tif image are stored in: <strong>Rigano et al._Figure 5_UseCase_AS_fswt-6hvirus-10minfix-stk_4-epi.tif.JSON</strong></p> <p><em><strong>Instructional video tutorials on how to use these example data files:</strong></em><br> Use these videos to get started with using Micro-Meta App after downloading the example data files available here.</p> <ul> <li><a href="https://vimeo.com/562022222">Part 1/2</a></li> <li><a href="https://vimeo.com/562022281">Part 2/2</a></li> </ul>
Metadata for the urbisphere-Paris campaign during 2022-2024: fieldwork maintenance log [L1]
<p>Machine-readable, formatted and redacted electronic fieldwork logs from the urbisphere-Paris observation campaign conducted between 2022-09-05 and 2024-07-22 in Paris, France. Provided in text format with comma separated columns (.csv) and in Microsoft Excel (.xlsx) format.</p> <p>The fieldwork logs are created from raw google form data submitted by campaign managers, scientists, technicinas and students. The formatting process is detailed in https://github.com/Urban-Meteorology-Reading/urbisphere-paris-fieldwork-log-format. The GitHub output has then been manually edited and adjusted.</p> <p>Contains maintenance information for the following observational sites operated as part of the urbisphere Paris campaign 2022 - 2024:</p> <table> <tbody> <tr> <td>PAARBO</td> <td>Paris – Arboretum de Vallée-aux-Loups </td> </tr> <tr> <td>PAAUNA</td> <td>Paris – Aunay-sous-Auneau</td> </tr> <tr> <td>PABOBI</td> <td>Paris – Bobigny</td> </tr> <tr> <td>PABONN</td> <td>Paris – Bonniel</td> </tr> <tr> <td>PABPAC</td> <td>Paris – Balloon Parc Andre Citroën</td> </tr> <tr> <td>PACHAM</td> <td>Paris – Chamant</td> </tr> <tr> <td>PACHAN</td> <td>Paris – Changis-sur-Marne </td> </tr> <tr> <td>PACHEM</td> <td>Paris – Chemin Vert Bobigny</td> </tr> <tr> <td>PACOMP</td> <td>Paris – Compiègne</td> </tr> <tr> <td>PACOUR</td> <td>Paris – Courdimanche-sur-Essonne</td> </tr> <tr> <td>PACRET</td> <td>Paris – Créteil</td> </tr> <tr> <td>PADENF</td> <td>Paris – Denfert Rocherau </td> </tr> <tr> <td>PADROU</td> <td>Paris – Droue Sur Drouette</td> </tr> <tr> <td>PAHOTE</td> <td>Paris – Hôtel de Ville</td> </tr> <tr> <td>PAJUSS</td> <td>Paris – Jussieu </td> </tr> <tr> <td>PALUPD</td> <td>Paris – Université Paris Diderot (LISA Platform)</td> </tr> <tr> <td>PAMEUD</td> <td>Paris – Meudon</td> </tr> <tr> <td>PANANG</td> <td>Paris – Nangis</td> </tr> <tr> <td>PANATI</td> <td>Paris – Rue Nationale</td> </tr> <tr> <td>PAPRUN</td> <td>Paris – Prunay-le-Temple</td> </tr> <tr> <td>PAROIS</td> <td>Paris – Roissy</td> </tr> <tr> <td>PAROMA</td> <td>Paris – Romainville</td> </tr> <tr> <td>PASIRT</td> <td>Paris – SIRTA Observatory Palaiseau</td> </tr> <tr> <td>PASTFE</td> <td>Paris – Saint Félix</td> </tr> <tr> <td>PAWYDT</td> <td>Paris – Wy-dit-Joli-Village</td> </tr> </tbody> </table> <p> </p>
SBC LTER: Reef: Kelp Forest Community Dynamics: Transect geospatial metadata
The data table provides summaries of depth information (mean, standard deviation and coefficient of variation) for all of the transects surveyed as part of the SBC LTER Kelp Forest Monitoring program. All data are expressed in meters, reference to mean lower low water (MLLW). Each value is the result of 160 observations, four at each meter (n=160). The sampling locations in this dataset are 40 meter transects at nine reef sites along the mainland coast of the Santa Barbara Channel and at two sites on the north side of Santa Cruz Island. The PDF document provides descriptive information of the transect such as headings and relative locations of the points along the transect. Data were recorded in 2010 and 2011, except that IVEE transect 3, 4, 5, 6, 7, and 8 were recorded in 2024.
Ansible Galaxy roles, versions, and metadata
<p>A dataset of Ansible roles accompanying the SCAM 2020 publication: R. Opdebeeck, A. Zerouali, C. Velázquez-Rodríguez, C. De Roover. “Does Infrastructure as Code Adhere to Semantic Versioning? An Analysis of Ansible Role Evolution”, In Proc. 20th Int. Working Conf. on Source Code Analysis and Manipulation, 2020.</p> <p><strong>Contents</strong></p> <p>- `repos.tar.gz`: Tar-ball of git repositories of all roles included in the dataset. The subdirectories in this archive are named using the role's Ansible Galaxy qualified name, i.e., `<namespace>.<role_name>`<br> - `roles.json`: A JSON file containing metadata extracted from Ansible Galaxy for each role.<br> - `repo_paths.json`: A mapping from Ansible Galaxy role IDs to their path in the `repos` directory.<br> - `tag_versions.json`: A mapping from Ansible Galaxy role IDs to the role's repository's git tags and metadata on these tags.<br> - `version_analysis.json`: Similar to `tag_versions.json`, but with additional filtering applied.<br> - `versiondiff_analysis.json`: Contains syntactical change statistics for each version bump in the role repositories.<br> - `structural_diff_analysis.json`: Contains structural change statistics for each version bump in the role repositories.<br> - `struct_diff_cache.zip`: Directory containing per-role diff statistics, primarily used for caching during the pipeline.<br> - `metrics_diffs_releases.csv`: CSV containing the structural diff statistics merged with the bump type of the version increments.<br> - `reports.zip`: Graphs and charts describing some of the output of the pipeline, as well as CSVs containing raw data.<br> - `version.json`: The version of the dataset structure.</p> <p><strong>Software and Tools</strong></p> <p>Tools to gather, extract, and process this data can be found separately at <a href="https://zenodo.org/record/4040647">https://zenodo.org/record/4040647</a>.</p>
Human intestinal Bacteria Collection (HiBC): Isolates and genomes metadata
<p>The <a href="https://hibc.rwth-aachen.de/" target="_blank" rel="noopener">Human intestinal Bacteria Collection (HiBC)</a> is a collection of bacterial strains, isolated from the human gut for which 16S rRNA gene sequences, genome sequences and culture conditions are made available to the research community. In addition to previously described bacteria, we include strains that represent novel species which have been taxonomically described and validly named, or will be in the future. This collection will be updated regularly.</p> <p>This dataset includes the taxonomy of the isolates, as well as metadata regarding their cultivation and isolation. We also provide metadata regarding the sequencing, genome assembly process and the biological sequences.</p> <p><strong>UPDATE v7</strong>: INSDC accession for <em>Segatella sinensis</em> CLA-AA-H117 was a missing value and is now the correct value of GCA_040324585.2.</p> <p><strong>UPDATE v6: </strong>The growth atmosphere is now indicated by anaerobic or aerobic instead of "Anaerobe/Aerobe" that was a misleading term. The risk group of these two isolates went from 1 to 2:</p> <ul> <li>CLA-AA-H205: <em>Anaerostipes caccae </em></li> <li>CLA-AA-H83: <em>Bacteroides fragilis</em></li> </ul> <p>The risk group of the following isolates has been updated (usually from unknown to 1, or from 2 to 1):</p> <ul> <li>CLA-SR-H026: <em>Aedoeadaptatus acetigenes</em></li> <li>CLA-KB-H139:<em> Bacteroides xylanisolvens</em></li> <li>CLA-SR-H015: <em>Bacteroides xylanisolvens</em></li> <li>CLA-AA-H187: <em>Blautia fusiformis</em></li> <li>CLA-AA-H274: <em>Brotaphodocola catenula</em></li> <li>CLA-AA-H286: <em>Butyricimonas faecihominis</em></li> <li>CLA-AA-H278:<em> Clostridium fessum</em></li> <li>CLA-AA-H147: <em>Dorea ammoniilytica</em></li> <li>CLA-SR-H027: D<em>orea formicigenerans</em></li> <li>CLA-KB-H89: <em>Dorea longicatena</em></li> <li>CLA-KB-H94: <em>Dorea longicatena</em></li> <li>CLA-SR-H022: <em>Enterococcus lactis</em></li> <li>CLA-AA-H250: <em>Hominenteromicrobium mulieris</em></li> <li>CLA-AA-H232: H<em>ominilimicola fabiformis</em></li> <li>CLA-AA-H246: <em>Hominisplanchenecus faecis</em></li> <li>CLA-AA-H276:<em> Hominiventricola filiformis</em></li> <li>CLA-AA-H213:<em> Oliverpabstia intestinalis</em></li> <li>CLA-AA-H241: <em>Oliverpabstia intestinalis</em></li> <li>CLA-AA-H58: <em>Pilosibacter fragilis</em></li> <li>CLA-KB-H110: <em>Ruthenibacterium lactatiformans</em></li> <li>CLA-AA-H174: <em>Segatella sinensis</em></li> <li>CLA-AA-H2: <em>Veillonella parvula</em></li> <li>CLA-AA-H273: <em>Waltera acetigignens</em></li> </ul> <p>Typos in media list have been fixed. </p> <p><strong>UPDATE v5</strong>: The accessions number for the genomes on INSDC databases are added under the column Accession. Plus two typos in the risk group column have been corrected as follow:</p> <ul> <li>CLA-AA-H173: from Risk Group 4 (!) to 2 like the other strain of <em>Sutterella wadsworthensis</em></li> <li>CLA-AA-H198: from Risk Group 4 (!) to 1 like the other <em>Bifidobacterium </em>species.</li> </ul> <p><strong>UPDATE v4</strong>: Only the taxonomy of a couple of isolates has been changed, as follow:</p> <ul> <li>CLA-ER-H4: <em>Collinsella sp900547855</em> instead of <em>Collinsella sp900544645</em></li> <li>CLA-AA-H142: <em>Pilosibacter fragilis</em> (<em>f__Clostridiaceae</em>) instead of <em>Sakamotonia hominis gen. nov.</em> (<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H58: <em>Pilosibacter fragilis </em>(<em>f__Clostridiaceae</em>) instead of <em>Sakamotonia hominis gen. nov. </em>(<em>f__Lachnospiraceae</em>)</li> <li>CLA-AA-H89B: <em>Lachnospira intestinalis sp. nov.</em> instead of <em>Lachnospira hominis sp. nov.</em></li> <li>CLA-JM-H10: <em>Lachnospira hominis sp. nov.</em> instead of <em>Lachnospira intestinalis sp. nov.</em></li> <li>CLA-JM-H7B: <em>Faecalibacterium taiwanense</em> instead of <em>Faecalibacterium faecis sp. nov.</em></li> <li>CLA-JM-H45: <em>Merdimmobilis hominis</em> instead of <em>Hominicola intestinalis gen. nov.</em></li> </ul> <p><strong>UPDATE v3</strong>: The genome of one of our isolate had been unfortunately swapped. This mistake has been now corrected on Zenodo and Coscine. The genome of <em>Segatella sinensis</em> CLA-AA-H117 should be considered correct with 103 contigs and 3 671 232 nt. Please note that the genome available at the NCBI is the correct one (GCA_040324585.2). Two typos regarding taxonomy have been corrected as well: <em>Maccoya intestinihominis</em> has been corrected to <em>Maccoyia intestinihominis</em> and <em>Faecousia faecis</em> to <em>Faecousia intestinalis</em>.</p>
OpenCitations Meta RDF dataset of agent roles metadata and its provenance information
<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>agent roles</strong> of bibliographic resources<strong> </strong>(<a href="http://purl.org/spar/pro/RoleInTime" target="_blank" rel="noopener">http://purl.org/spar/pro/RoleInTime</a>). These agents can be authors, editors, or publishers. It contains all the metadata and its provenance information, structured specifically around agent roles, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /ar/06250/10000/1000/1000.zip, while information about provenance in /ar/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
OpenCitations Meta RDF dataset of page numbers metadata and its provenance information
<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>page numbers</strong> of bibliographic resources, known as <strong>manifestations </strong>(<a href="http://purl.org/spar/fabio/Manifestation" target="_new">http://purl.org/spar/fabio/Manifestation</a>). It contains all the bibliographic metadata and its provenance information, structured specifically around manifestations (page numbers), in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
OpenCitations Meta RDF dataset of identifiers metadata and its provenance information
<p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>identifiers </strong>(<a href="http://purl.org/spar/datacite/Identifier" target="_blank" rel="noopener">http://purl.org/spar/datacite/Identifier</a>) of bibliographic resources. It contains all the metadata and its provenance information, structured specifically around identifiers, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /id/06250/10000/1000/1000.zip, while information about provenance in /id/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p>
OpenCitations Meta RDF dataset of bibliographic resources metadata and its provenance information
<div> <p>This dataset is a specialized subset of the OpenCitations Meta RDF data, focusing exclusively on data related to <strong>bibliographic resources </strong>(<a href="http://purl.org/spar/fabio/Expression" target="_blank" rel="noopener">http:///purl.org/spar/fabio/Expression</a>). It contains all the metadata and its provenance information, structured specifically around bibliographic resources, in JSON-LD format.</p> <p>The inner folders are named through the <strong>supplier prefix</strong> of the contained entities. It is a prefix that allows you to recognize the entity membership index (e.g., OpenCitations Meta corresponds to <strong>06*0</strong>).</p> <p>After that, the folders have <strong>numeric names</strong>, which refer to the range of contained entities. For example, the 10000 folder contains entities from 1 to 10000. Inside, you can find the <strong>zipped </strong>RDF data.</p> <p>At the same level, additional folders containing the <strong>provenance </strong>are named with the same criteria already seen. Then, the 1000 folder includes the provenance of the entities from 1 to 1000. The provenance is located inside a folder called <strong>prov</strong>, also in zipped JSON-LD format.</p> <p>For example, data related to the entity is located in the folder /br/06250/10000/1000/1000.zip, while information about provenance in /br/06250/10000/1000/prov/se.zip</p> <p>Additional information about OpenCitations Meta at the <a href="https://opencitations.net/meta" target="_blank" rel="noopener">official webpage</a>.</p> <p> </p> </div>
The Semantic Turkey metadata registry ontology
<p>An application profile of DCAT combining it with other metadata vocabularies (e.g. VoID, DCTERMS, LIME) to meet requirements elicited in various use cases of the Semantic Web platform Semantic Turkey</p>
DZD Core Data Set - Metadata and SOP
<p>DZD Core Data Set - Metadata and SOPs contains documentation for the DZD Core Data Set. <a href="https://www.dzd-ev.de/en/">The German Center for Diabetes Research (DZD)</a> conducts large <a href="https://www.dzd-ev.de/en/research/multicenter-studies/index.html">clinical multicenter studies</a> in the field of diabetes and metabolic research. The DZD has established the DZD Core Data Set which contains a list of clinical parameters relevant for diabetes research in related clinical studies. The Core Data Set itself is published at MDM Portal.</p>
MALDI MS data and metadata from "A biocodicological analysis of the medieval library and archive from Orval Abbey, Belgium"
<p>See <a href="https://doi.org/10.1098/rsos.210210">Ruffini-Ronzani et al</a>.</p>
MetFrag Local CSV: CompTox (7 March 2019 release) MetaData File
<p>These are the CSV files that can be used as a local database in MetFrag (https://msbi.ipb-halle.de/MetFrag/), for those who wish to integrate this into the command line version.</p> <p>Note that this file is TOO LARGE to be uploaded via the web interface, updated versions of these files are integrated in the web interface.</p> <p>This upload includes the SelectMetaData version of the CompTox MetFrag file from the 7 March 2019 release (DOI:<a href="https://doi.org/10.23645/epacomptox.7525199.v2">10.23645/epacomptox.7525199.v2</a>).</p>
Polifonia Corpus - Encyclopedic Module Metadata - Spanish Language
<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at https://github.com/polifonia-project/Polifonia-Corpus</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.