Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

404

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

404 results for “reference data”

Learn how ShareScore rates datasets ↗
zenodo44/100

Data for publication "Benefits of open access to researchers from lower-income countries: A global analysis of reference patterns in 1980–2020"

<p>Data to reproduce figures for the publication "Benefits of open access to researchers from lower-income countries: A global analysis of reference patterns in 1980&ndash;2020" (DOI: 10.1177/01655515241245952). Each file contains the data underlying the figure corresponding to the file name.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

R package n2khab: providing preprocessed reference data for Flemish Natura 2000 habitat analyses

The n2khab package is an R package with preprocessing functions and standard reference data, useful for analyses regarding Flemish Natura 2000 habitats and regionally important biotopes (RIBs). URL: <a href="https://inbo.github.io/n2khab">https://inbo.github.io/n2khab</a>.

opengpl-3.0Jan 2025View details →
zenodo44/100

Raw and aggregated data for the study introduced in the paper "The way we cite: common metadata used across disciplines for defining bibliographic references"

<p>These data have been gathered in the context of a study aiming to investigate citation practices for referencing different types of entities and, in particular, for understanding the most used metadata in bibliographic references. The data are stored in two documents in XLSX format:</p> <ul> <li>file &quot;links-intext-pointers-and-cited-entity-types.xlsx&quot; - it contains information about whether the in-text reference pointers of the various PDF articles of the corpus have specified hypertextual links from the in-text reference pointers to the denoted bibliographic reference, plus information about the types of all the entities cited by each article in the corpus;</li> <li>file &quot;metadata-bibliographic-references.xlsm&quot; - it contains information about the metadata used to identify the various descriptive elements of all the bibliographic references defined in the article of the corpus.</li> </ul> <p>The methodology used to gather all these data is described in:</p> <blockquote> <p>Santos, E. A. d., Peroni, S., Mucheroni, M. L.: Workflow for retrieving all the data of the analysis introduced in the article &quot;Citing and referencing habits in Medicine and Social Sciences journals in 2019&quot;. (2020), <a href="https://doi.org/10.17504/protocols.io.bbifikbn">https://doi.org/10.17504/protocols.io.bbifikbn</a></p> </blockquote>

opencc-by-4.0May 2022View details →
zenodo44/100

Zonal mean data of the SOCOLv4 reference experiment (1980-2018)

<p>Zonal mean data of the SOCOLv4 reference experiment (1980-2018). O3, NOx, ClOx, HOx, BrOx, Temperature.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Data used in ECLIPSER methods paper and GTEx snRNA-seq cross-tissue reference map analysis

<p>The tables were used in the papers: Rouhana*, Wang* <em>et al.,</em>&nbsp;ECLIPSER: identifying causal cell types and genes for complex traits through single cell enrichment of e/sQTL-mapped genes in GWAS loci, bioRxiv 2021, doi: https://doi.org/10.1101/2021.11.24.469720; and Eraslan&nbsp;<em>et al.,</em>&nbsp;Single-nucleus cross-tissue molecular reference maps to decipher disease gene function, bioRxiv 2021,&nbsp;doi: https://doi.org/10.1101/2021.07.19.452954. &#39;<a href="https://zenodo.org/api/files/1f8d48d0-6bf7-4bec-b6ef-5c7a9ead8079/GTEx_v8_HG38_all_variants.tsv.gz">GTEx_v8_HG38_all_variants.tsv.gz</a>&#39; is&nbsp;an input file for running GWASvar2gene on GTEx v8 eQTLs and sQTLs, and all other files are input files for&nbsp;ECLIPSER.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

MALDI-TOF-MS reference spectra and sequence data for domesticated equids (horse and donkey) collagen for Zooarchaeology by Mass Spectrometry (ZooMS)

<p>MALDI-TOF-MS spectra of extracted collagen from modern reference and archaeological bone samples to develop markers for Zooarchaeology by Mass Spectrometry (ZooMS) to distinguish between Equus species. &nbsp;For each sample digestions were done in both trypsin and chymotrypsin separately. &nbsp;Information about the species of the samples can be found in &#39;sample metadata.csv&#39; file. &nbsp;Information on the extraction and digestion protocol can be found in the associated manuscript. The sequence data contains alignments of the proteins COL1A1 and COL1A2 for available Equus collagen protein sequences. &nbsp;More information on these files can be found in the corresponding manuscript to this dataset.<br> &nbsp;</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Control panel created from 30 Nanopore sequencing data from the Human Pangenome Reference Consortium

<p>This is control panel for <a href="https://github.com/friend1ws/nanomonsv">nanomonsv</a> software, which is expected to exclude many false positives as well as improve computational cost. This is made by aligning 30 Nanopore sequencing data from Human Pangenome Reference Consortium to the GRCh38 reference genome (obtained from&nbsp;<a href="https://console.cloud.google.com/storage/browser/genomics-public-data/resources/broad/hg38/v0;tab=objects">here</a>) with <a href="https://github.com/lh3/minimap2">minimap2</a> version 2.24.&nbsp;<strong>When you use these control panels and publish, do not forget to credit to&nbsp;<a href="https://humanpangenome.org/data-use-protocol/">HPRC</a>!</strong></p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Comprehensive 100-bp resolution genome-wide epigenomic profiling data for the hg38 human reference genome

<p>This is a comprehensive collection of diverse epigenomic profiling data in 100-bp resolution with full genome-wide coverage. The datasets are processed from raw read count data collected from five types of sequencing-based assays collected by the Encyclopedia of DNA Elements (ENCODE, <a href="http://www.encodeproject.org">http://www.encodeproject.org</a>) consortium. A total of 6,305 alignment profiles from various high-throughput sequencing assays available on the ENCODE database were preprocessed and filtered according to ENCODE&rsquo;s data standard</p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

Reference Data Set: Electricity, Heat, and Gas Sector Data for Modeling the German System

<p>This reference data set representing the status quo of the German electricity, heat, and natural gas sectors was compiled within the research project &lsquo;LKD-EU&rsquo; (Long-term planning and short-term optimization of the German electricity system within the European framework: Further development of methods and models to analyze the electricity system including the heat and gas sector).</p> <p>While the focus is on the electricity sector, the heat and natural gas sectors are covered as well. With this reference data set, we aim to increase the transparency of energy infrastructure data in Germany. Where not otherwise stated, the data included in this report is given with reference to the year 2015 for Germany. The data set is documented in DIW Data Documentation 92 (see references).</p> <p>The project is a joined effort by the German Institute for Economic Research (DIW Berlin), the Workgroup for Infrastructure Policy (WIP) at Technische Universit&auml;t Berlin (TUB), the Chair of Energy Economics (EE2) at Technische Universit&auml;t Dresden (TUD), and the House of Energy Markets &amp; Finance at University of Duisburg-Essen. The project was funded by the German Federal Ministry for Economic Affairs and Energy through the grant &lsquo;LKD-EU&rsquo;, FKZ 03ET4028A-D.</p>

openother-openDec 2017View details →
zenodo44/100

Preliminary data collected by 2 prototype Surface Velocity Platform drifters with Barometer and Reference Sensor for Temperature (SVP-BRST)

<p>The SVP-BRST drifter was developed to serve calibration and validation of Sentinel satellite SST retrievals. Two prototypes were deployed in the Mediterranean Sea end of April 2018. Preliminary data collected then until 11 June 2018 are published in this dataset. The drifters were developed and deployed under funding from the European Union&#39;s Copernicus Programme. The data are transmitted from the buoy to shore using data format #091 (see References).</p>

opencc-by-4.0Sep 2018View details →
zenodo44/100

Reference data for the Perram and Wertheim (1985) contact function of ellipsoids

<p>Reference data for the Perram and Wertheim (1985) contact function of ellipsoids</p> <p>This dataset provides reference values of the contact function of two ellipsoids, as defined by Perram and Wertheim (Perram, J. W., &amp; Wertheim, M. S. (1985). Statistical mechanics of hard ellipsoids. I. Overlap algorithm and the contact function. Journal of Computational Physics, 58(3), 409&ndash;416. <a href="https://doi.org/10.1016/0021-9991(85)90171-8">DOI:10.1016/0021-9991(85)90171-8</a>). This paper will be referred to as PW85 in what follows.</p> <p>Reference values of the <code>F</code> function</p> <p>The data is shared as a HDF5 file <code>pw85_ref_data-YYYYMMDD.h5</code>, which contains the following datasets (to be described below)</p> <ul> <li><code>directions</code>: a 12&times;3 array,</li> <li><code>F</code>: a 108&times;108&times;12&times;9 array,</li> <li><code>lambdas</code>: a length-9 array,</li> <li><code>radii</code>: a length-3 array,</li> <li><code>spheroids</code>: a 108&times;6 array.</li> </ul> <p>The attached Python script <code>pw85_gen_ref_data.py</code> was used to generate the data; it uses the <a href="http://mpmath.org/">mpmath</a> library.</p> <p>Mathematical definition of the contact function</p> <p>The contact function is defined in PW85 as the maximum over <code>(0, 1)</code> of the <code>F</code> function which is defined as follows [see Eq. (3.7) in PW85, with slightly different notations]</p> <pre><code>F(λ) = λ(1-λ)r₁₂ᵀ⋅Q⁻¹⋅r₁₂,</code></pre> <p>where <code>0 &le; &lambda; &le; 1</code> is a scalar, <code>r₁₂</code> is the center-to-center vector. <code>Q</code> is the matrix defined as follows</p> <pre><code>Q = (1-λ)Q₁ + λQ₂,</code></pre> <p>where <code>Qᵢ</code> is the symmetric, positive definite matrix that defines ellipsoid <code>&Omega;ᵢ</code> through</p> <pre><code>m ∈ Ωᵢ iff (m-cᵢ)ᵀ⋅Qᵢ⁻¹⋅(m-cᵢ) ≤ 1,</code></pre> <p>where <code>cᵢ</code> is the center of <code>&Omega;ᵢ</code>. Then, the contact function <code>F₁₂</code> is defined as the maximum of <code>F</code> [see Eq. (3.8) in PW85]</p> <pre><code>F₁₂(r₁₂, Q₁, Q₂) = max{ F(λ), 0 ≤ λ ≤ 1 }.</code></pre> <p>Parametrization</p> <p>The reference data is restricted to spheroids (equatorial radius: <code>aᵢ</code>; polar radius: <code>cᵢ</code>; direction of axis of revolution: <code>nᵢ</code>)</p> <pre><code>Qᵢ = aᵢ²I + (cᵢ²-aᵢ²)nᵢᵀ⋅nᵢ,</code></pre> <p>(<code>I</code>: identity matrix). The radii take the following values</p> <pre><code>aᵢ, cᵢ ∈ {0.01999, 1.999, 9.999}.</code></pre> <p>These values of the radii are stored in the <code>radii</code> dataset of the HDF5 file. The orientations <code>nᵢ</code> coincide with the vertices of an icosahedron</p> <pre><code>nᵢ = [0, ±u, ±v]ᵀ or nᵢ = [±v, 0, ±u]ᵀ or nᵢ = [±u, ±v, 0]ᵀ,</code></pre> <p>where</p> <pre><code> 1 φ 1+√5 u = ───────, v = ─────── and φ = ────. √(1+φ²) √(1+φ²) 2</code></pre> <p>The orientations are stored in the <code>directions</code> dataset as a 12&times;3 array. The matrices <code>Qᵢ</code> are precomputed and stored in the <code>spheroids</code> dataset as a 108&times;6 array (note: 108 = 12 orientations &times; 3 equatorial radii &times; 3 polar radii). <code>spheroids[i, :]</code> stores the upper triangular part of the corresponding matrix in row-major order</p> <pre><code>⎡ spheroids[i, 0] spheroids[i, 1] spheroids[i, 2] ⎤ ⎢ spheroids[i, 3] spheroids[i, 4] ⎥. ⎣ sym. spheroids[i, 5] ⎦</code></pre> <p>The scalar <code>&lambda;</code> takes tabulated values (see the <code>lambdas</code> dataset)</p> <pre><code>λ ∈ {0.1, 0.2, …, 0.9}.</code></pre> <p>Note that <code>&lambda; = 0.0</code> and <code>&lambda; = 1.0</code> are excluded, since <code>F</code> is uniformly 0 in that case.</p> <p>Reference values of the <code>F</code> function</p> <p>The reference values of the function <code>F</code> are stored in the <code>F</code> dataset, which is a 108&times;108&times;12&times;9, such that <code>F[i, j, h, k]</code> is the value of <code>F</code> for</p> <pre><code>Q₁ = spheroids[i], Q₂ = spheroids[j], r₁₂ = directions[h] and λ = lambdas[k].</code></pre> <p>Note that the <code>r₁₂</code> vector takes values in the <code>directions</code> dataset. In other words, only unit-length center-to-center vectors are considered here. Indeed, <code>F</code> trivially depends on the norm of <code>r₁₂</code>, which is therefore not considered here in order to reduce the size of the dataset.</p> <p>Reference values of the contact function</p> <p>Note: the following is <em>not</em> implemented yet, as reference values of the contact function were not deemed useful. Indeed, once <code>F</code> is validated, it is straightforward to check that the implementation of <code>F₁₂</code> to be tested indeed maximizes <code>F</code>.</p> <blockquote> <p>The reference values of the contact function <code>F₁₂</code> are stored in the <code>contact_function</code> dataset, which is a 108&times;108&times;12&times;3 array, such that <code>contact_function[i, j, h, k]</code> is the value of <code>F₁₂</code> for</p> <pre><code>Q₁ = spheroids[i], Q₂ = spheroids[j] and r₁₂ = radii[h] * directions[k].</code></pre> <p>Note that the <code>r₁₂</code> vector is not normed, here.</p> </blockquote>

opencc-by-4.0Jul 2019View details →
zenodo44/100

AirHeritage Datalake: Multi-site, Multi-season, Multi Unit dataset including Fixed and Mobile Citizen science data from networked Air Quality Low-Cost Multi-Sensors devices and reference stations

<p>This datalake comprises several datasets from <strong>37 networked low cost air quality multisensors</strong> (<strong>30</strong> <strong>mobile</strong> ENEA MONICA(tm) +&nbsp;<strong>7</strong> <strong>fixed</strong>) along with <strong>3</strong> (fixed) + <strong>1</strong> (mobile) <strong>reference stations</strong> operated by Campania Regional Envronmental Protection Agency. The datalake is organized in 3 main directories respectively related to fixed nodes, mobile nodes and nearby reference stations including a mobile laboratory used for colocation campaigns; each subdirectory include its own metadata description file.</p> <p>Data, curated by Energy and Data Science Laboratory of ENEA, include multi-weeks colocation periods when low cost devices have been colocated with reference stations as well as operational periods during which sensors are deployed for fixed or mobile monitoring campaigns. Data have been recorded during 2021 and 2022 in a<strong> pervasive, multi-site, multi-seasonal deployment</strong> in Portici, a densely populated small area city (4km2, 55k + inhabitants) located 7km south of Naples, Italy.</p> <p>The datalake consists in actual sensors and reference intrumentations timeseries along with metadata description files with&nbsp; &nbsp;deployment dates and location data. The dataset files include high sampling frequency raw sensor data of quality-controlled sensor network along with co-located reference stations data sets. Sensor data include electrochemical sensors data (intended target pollutants: NO2, O3, CO), Optical sensor data (PM2.5, PM10, PM1) readings along with meteorological parameters. .</p> <p>Further description of sensors and reference instruments are reported in the accompanying paper (see citation request).</p> <p>The dataset can be used for&nbsp;</p> <ul> <li>&nbsp;<strong>advanced (remote/universal/in field) data driven calibration strategies</strong> test or development including <strong>machine learning </strong>models</li> <li><strong>mobile opportunistic data fusion</strong> methods development</li> <li><strong>geomatics and data assimilation</strong> models studies</li> </ul> <p>as well as low cost sensor characterization performance studies.&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types probabilities (part 1)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types probabilities (part 2)</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Soil type (World Reference Base) maps of Europe based on Ensemble Machine Learning and multiscale EO data

<h2><strong>Sub-dataset: WRB soil types classification and relative entropy</strong></h2> <h2>Disclaimer</h2> <p>This is the first release of pan-EU predictions of soil health indicators (the Soil Health Data Cube). Use for testing purposes only. A publication describing methods used has been submitted to PeerJ and is in review. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or European Commision. Neither the European Union nor the granting authority can be held responsible for them. The data is provided "as is". AI4SoilHealth project consortium and its suppliers and licensors hereby disclaim all warranties of any kind, express or implied, including, without limitation, the warranties of merchantability, fitness for a particular purpose and non-infringement. Neither AI4SoilHealth project Consortium nor its suppliers and licensors, makes any warranty that the Website will be error free or that access thereto will be continuous or uninterrupted. You understand that you download from, or otherwise obtain content or services through, the Website at your own discretion and risk.</p> <h2>Description</h2> <p>This dataset covers pan-European areas, including Ukraine, the UK, and Turkey. This data cube could be used for applications such as soil property mapping and comprehensive soil health assessment across Europe. The dataset spans four depth ranges and multiple time periods, providing information for studies on soil organic carbon stock and dynamics.</p> <p>This dataset is part of the "Soil type (World Reference Base) map of Europe based on Ensemble Machine Learning and multiscale EO data" dataset. Check the related identifiers section below to access other parts of the dataset.</p> <p>This data set includes:</p> <ul> <li><strong>Soil types classification and relative entropy:</strong><br> This data includes hard classes maps (185 soil type classes) produced by ensemble model and relative entropy (Kullback-Leibler divergence) maps in added information (bit) over a dummy distribution (scaled 1000x). </li> <li><strong>Soil types probabilities (part 1):</strong><br> This data includes 92 averaged probabilities (0-1) maps for classes from <strong>abruptic.acrisols</strong> to <strong>gleyic.arenosols</strong>. The probabilites were scaled 100x (0-100). </li> <li><strong>Soil types probabilities (part 2):</strong><br> This data includes 93 averaged probabilities maps for classes from <strong>gleyic.cambisols</strong> to <strong>vitric.andosols</strong>. The probabilites were scaled 100x (0-100). </li> </ul> <h3>Related identifiers</h3> <ul> <li><a href="https://zenodo.org/records/13838407">WRB soil types classification and relative entropy</a></li> <li><a href="https://zenodo.org/records/13837830">WRB soil types probabilities (part 1)</a></li> <li><a href="https://zenodo.org/records/13837832">WRB soil types probabilities (part 2)</a></li> </ul> <h3>Data Details</h3> <ul> <li><strong>Time period:</strong> long term.</li> <li><strong>Type of data:</strong> Soil types classification and model probabilities.</li> <li><strong>How the data was collected or derived:</strong> The data was derived using ensemble ML models.</li> <li><strong>Statistical methods used:</strong> Relative entropy (Kullback-Leibler divergence)</li> <li><strong>Limitations or exclusions in the data:</strong> The dataset does not include data for Svalbard. </li> <li><strong>Coordinate reference system:</strong> EPSG:3035</li> <li><strong>Bounding box (Xmin, Ymin, Xmax, Ymax):</strong> (900,000, 899,000, 7,401,000, 5,501,000)</li> <li><strong>Spatial resolution:</strong> 30m</li> <li><strong>Image size:</strong> 216,700P x 153,400L</li> <li><strong>File format:</strong> Cloud Optimized Geotiff (COG) format.</li> </ul> <h3>Support</h3> <p>If you discover a bug, artifact, or inconsistency, or if you have a question please raise a GitHub issue: GitLab Issues (tbc)</p> <h3>Name convention</h3> <p>To ensure consistency and ease of use across and within the projects, we follow the standard Ai4SoilHealth and Open-Earth-Monitor file-naming convention. The convention works with 10 fields that describe important properties of the data. In this way users can search files, prepare data analysis etc, without needing to open files. For example, in <strong>soil.types_ai4sh.ensemble_c_30m_s_20220101_20221231_epsg.3035_v20240917.tif</strong>, the fields are:</p> <ol> <li><strong>generic variable name:</strong> soil.types = soil types</li> <li><strong>variable procedure combination:</strong> ai4sh.ensemble.abruptic.acrisols = AI4SH project, ensemble model, abrupitc acrisols soil type.</li> <li><strong>Position in the probability distribution/variable type:</strong> m = mean | c = class | p = probability</li> <li><strong>Spatial support:</strong> 30m</li> <li><strong>Depth reference:</strong> s = surface</li> <li><strong>Time reference begin time:</strong> 20220101 = 2022-01-01</li> <li><strong>Time reference end time:</strong> 20221231 = 2022-12-31</li> <li><strong>Bounding box:</strong> eu = pan-Europe</li> <li><strong>EPSG code:</strong> epsg.3035</li> <li><strong>Version code:</strong> v20240917 = version from 2024-09-17</li> </ol>

opencc-by-4.0Sep 2024View details →
zenodo44/100

RE-Lab-Projects/TRY_DE_2015_2045: Test Reference Years (TRY) for 15 typical regions in germany with special regards on realisitc radiation data on a 1min timescale

<p>Test Reference Years (TRY) for 15 typical regions in germany with special regards on realisitc radiation data on a 1min timescale</p> <p><strong>Summary:</strong></p> <p>The data set contains the updated test reference years (TRY) of the German Weather Service (DWD). By subdividing into 15 TRY regions, each postcode area can be assigned a representative weather data set. It should be emphasized that in addition to a mean, current test reference year for a region, there is also a year with extreme summer and extreme winter weather. To take climate change into account, there is then a time series for the year 2045 for each test reference year based on the IPCC climate models. This means that a total of 90 weather data sets are available with a one-hour time resolution.</p> <p>In order to use the data in simulations with a temporal resolution of 1min or 15min, the data set was extended by linear interpolation. While this approach is justifiable for air pressure and temperature, for example, it does not depict high fluctuations in solar radiation. Therefore, based on the one-minute open data measurement data set of the Baseline Surface Radiation Network, with an algorithm by Hofmann et. al. the time series of global radiation are newly generated for all test reference years. Another algorithm by Hofmann et. al. was used to calculate the corresponding diffuse radiation times series.</p> <p><strong>Sources:</strong></p> <ul> <li>Raw data from DWD: <a href="https://kunden.dwd.de/obt/">https://kunden.dwd.de/obt/</a> -&gt; <code>1_raw-data</code></li> <li>Synthetic 1min radiation data: <a href="http://pvmodelling.org/">http://pvmodelling.org/</a> -&gt; <code>2_synthetic-radiation</code></li> </ul> <p><strong>How to use or recreate the final dataset:</strong></p> <ol> <li>clone/download this repository</li> <li>unzip the files from the data.zip file <ol> <li><a href="https://github.com/RE-Lab-Projects/TRY_DE_2015_2045/releases/download/v1.4.0/data.zip">https://github.com/RE-Lab-Projects/TRY_DE_2015_2045/releases/download/v1.4.0/data.zip</a></li> </ol> </li> <li>Use or recreate the final dataset <ol> <li>use: Final datasets are then located in -&gt; <code>3_processed-data</code></li> <li>recreate: run the <code>process-data.py</code></li> </ol> </li> </ol> <p><strong>Test reference stations / regions</strong></p> <p>No. | lon | lat | station | region<br> 1 | 53.5591 | 8.5872 | Bremerhaven | Nordseek&uuml;ste<br> 2 | 54.0878 | 12.1088 | Rostock | Ostseek&uuml;ste<br> 3 | 53.5299 | 10.0078 | Hamburg | Nordwestdeutsches Tiefland<br> 4 | 52.3938 | 13.0651 | Potsdam | Nordostdeutsches Tiefland<br> 5 | 51.4562 | 7.0568 | Essen | Niederrheinisch-westf&auml;lische Bucht und Emsland<br> 6 | 550.6461 | 7.9426 | Bad Marienburg | N&ouml;rdliche und westliche Mittelgebirge, Randgebiete<br> 7 | 51.3334 | 9.4725 | Kassel | N&ouml;rdliche und westliche Mittelgebirge, zentrale Bereiche<br> 8 | 51.7239 | 10.6069 | Braunlage | Oberharz und Schwarzwald (mittlere Lagen)<br> 9 | 50.8233 | 12.9181 | Chemnitz | Th&uuml;ringer Becken und S&auml;chsisches H&uuml;gelland<br> 10 | 50.3226 | 11.9124 | Hof | S&uuml;d&ouml;stliche Mittelgebirge bis 1000 m<br> 11 | 50.4312 | 12.9522 | Fichtelberg | Erzgebirge, B&ouml;hmer- und Schwarzwald oberhalb 1000 m<br> 12 | 49.4902 | 8.4637 | Mannheim | Oberrheingraben und unteres Neckartal<br> 13 | 48.2432 | 12.5286 | M&uuml;hldorf | Schw&auml;bisch-fr&auml;nkisches Stufenland und Alpenvorland<br> 14 | 48.6536 | 9.8666 | St&ouml;tten | Schw&auml;bische Alb und Baar<br> 15 | 47.4945 | 11.1046 | Garmisch Partenkirchen | Alpenrand und -t&auml;ler</p> <p><strong>Content</strong></p> <ul> <li><strong>files</strong>: 90 test reference years (TRY) <pre><code>15 test reference regions x 3 reference conditions (average year, extreme summer, extreme winter) x 2 reference projections (year 2015 and year 2045) </code></pre> </li> <li><strong>columns per file</strong>: <pre><code>datetime [yyyy-MM-dd hh:mm:ss+01:00/02:00] temperature [degC] pressure [hPa] wind direction [deg] wind speed [m/s] cloud coverage [1/8] humidity [%] direct irradiance [W/m^2] diffuse irradiance [W/m^2] synthetic global irradiance [W/m^2] synthetic diffuse irradiance [W/m^2] clear sky irradiance [W/m^2] </code></pre> </li> <li><strong>length</strong>: 1 year</li> <li><strong>time increment</strong>: 60s / 900s / 3600s</li> </ul> <p><strong>Important hints</strong>:</p> <ul> <li>all files in <code>3_processed-data</code> were calculated with the skript <code>process-data.py</code></li> <li><em>A value with, for example, a timestamp 12:00:00 represents the mean value from this timestamp until the following timestamp.</em></li> <li><em>datetime column is in CET / CEST</em></li> </ul>

opencc-by-4.0Jul 2021View details →
zenodo44/100

SEMAFORA Semantic Reference Data Models

<p>To support the aim of the Semafora project, a series of Semantic Reference Data Models were created to provide a target semantic structure for the integration of standard archaeological survey data.&nbsp;</p> <p>&nbsp;</p> <p>The following models constitute the Semafora SRDM package:</p> <p>&nbsp;</p> <p>Place: This model is used to document any places associated with the archaeological survey.</p> <p>&nbsp;</p> <p>Institution: This model is used to document any institution associated with the survey.</p> <p>&nbsp;</p> <p>Period: This model is used to document the generic historical period assigned to the production of artefacts, existence of sites or other observable archaeological and historical events.</p> <p>&nbsp;</p> <p>Feature: This model is used to document any physical features, such as walls and other human-made structures observable on the field.</p> <p>&nbsp;</p> <p>Project: This model is used to document the overarching project, a part of which is the archaeological survey. Some projects may involve surveys, excavations, and other archaeological activities.</p> <p>&nbsp;</p> <p>Site: This model is used to document a site declared as archaeological as a result of the survey process.</p> <p>&nbsp;</p> <p>Digital Object: This model is used to document any type of digital asset associated with the survey.</p> <p>&nbsp;</p> <p>Survey Unit: This model is used to document a defined survey unit where the survey activity happens. It has both the properties of a place with dimensions and coordinates and of a physical thing from which samples can be collected.</p> <p>&nbsp;</p> <p>Collection: This model is used to document a collection of physical things, usually artefacts, collected while surveying.</p> <p>&nbsp;</p> <p>Artefact: This model is used to document individual artifacts collected from while surveying as a part of a larger collection of material things or as a singular artefact collection or documentation.</p> <p>&nbsp;</p> <p>Image: This model is used to document any image representing components of the archaeological survey, such as artefacts, features, places, people, etc.</p> <p>&nbsp;</p> <p>Observation: This model is used to document the act of observation usually associated with archaeological sites or survey units and the properties assigned to those as a result of the observation.</p> <p>&nbsp;</p> <p>Bibliography: This model is used to document any textual object associated with the survey or any components of it.</p> <p>&nbsp;</p> <p>Sample: This model is used to document a material sample of the survey unit. It partially overlaps with collection but acts as a parent sample that may contain other physical things besides human-made objects.</p> <p>&nbsp;</p> <p>Person: This model is used to document an individual person (alive or dead) involved in some way in the survey process.</p> <p>&nbsp;</p> <p>These models are intended to be used in order to guide semantic data mapping processes as well as to provide instructions for the creation of a target data semantic data management system.</p> <p>&nbsp;</p> <p>Each model&rsquo;s semantic reference data model description is stored here as a csv. The ongoing curation and updating of these SRDMs is undertaken using the Zellij system and can be accessed here:</p> <p>&nbsp;</p> <p><a href="https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c">https://zellij.pythonanywhere.com/docs/list/appaCfYmUmH78z85c</a></p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Reference data set for a Norwegian medium voltage power distribution system

<p>This reference data set describes a representative Norwegian radial, medium voltage (MV) electric power distribution system operated at 22 kV. The data set is developed in the Norwegian research centre CINELDI and will in brief be referred to as the CINELDI MV reference system.</p> <p>Data for a real Norwegian distribution system were provided by a distribution grid company. The data have been anonymized and processed to obtain a simplified but still realistic grid model with 124 nodes. The data set consists of the following three parts:<br> 1. Grid data files: describe the base version of the reference system that represents the present-day state of the grid, including information about topology, electrical parameters, and existing load points.<br> 2. Load data files: comprise load demand time series for a year with hourly resolution and scenarios for the possible long-term development of peak load. These data describe an extended version of the reference system with information about possible new load points being added to the system in the future.<br> 3. Reliability data files: contain data necessary for carrying out reliability of supply analyses for the system.</p> <p>The data set is described in detail in the following data article:<br> I. B. Sperstad, O. B. Fosso, S. H. Jakobsen, A. O. Eggen, J. H. Evenstuen, and G. Kj&oslash;lle, &ldquo;Reference data set for a Norwegian medium voltage power distribution system,&rdquo; Data in Brief, 109025, 2023, doi: 10.1016/j.dib.2023.109025.</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Data of the characterisation of conventional 87Sr/86Sr isotope ratios in cement, limestone and slate reference materials based on an interlaboratory comparison study

<p>This dataset represents the electronic supplementary material (ESM) of the publication entitled &quot;Characterisation of conventional <sup>87</sup>Sr/<sup>86</sup>Sr isotope ratios in cement, limestone and slate reference materials based on an interlaboratory comparison study&quot;, which is published in Geostandards and Geoanalytical Research under the DOI: 10.1111/GGR.12517. It consists of four files. &#39;ESM_Data.xlsx&#39; contains all reported data of the participants, a description of the applied analytical procedures, basic calculations, the consensus values, and part of the uncertainty assessment. &#39;ESM_Figure-S1&#39; displays a schematic on how measurements, sequences and replicates are treated for the uncertainty calculation carried out by PTB. &#39;ESM_Technical-protocol.pdf&#39; is the technical protocol of the interlaboratory comparison, which has been provided to all participants together with the samples and which contains bedside others the definition of the measurand and guidelines for data assessment and calculations. &#39;ESM_Reporting-template.xlsx&#39; is the Excel template which has been submitted to all participants for reporting their results within the interlaboratory comparison. Excel files with names of the the structure &#39;GeoReM_Material_Sr8786_Date.xlsx&#39; represent the <em>R</em><sub>con</sub>(<sup>87</sup>Sr/<sup>86</sup>Sr) data for a specific reference material downloaded from GeoReM at the specified date, e.g. &#39;GeoReM_IAPSO_Sr8786_20221115.xlsx&#39; contains all <em>R</em><sub>con</sub>(<sup>87</sup>Sr/<sup>86</sup>Sr) data for the IAPSO seawater standard listed in GeoReM until 15 November 2022.</p>

opencc-by-4.0Apr 2023View details →
edi44/100

End of growing season percent cover and biomass data for hayed vs reference sites, Rowley, MA, PIE LTER

End of growing season percent cover and biomass data of marsh vegetation in hayed and reference marsh sites near Stackyard Rd. and Patmos Rd, Rowley, Massachusetts.

openCC (other)Jan 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record