Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

8,038

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

8,038 results for “validation”

Learn how ShareScore rates datasets ↗
zenodo44/100

OpenFOAM cases for the Validation of the CHT Model

<p>This dataset contains the<em>&nbsp;underlying data</em>&nbsp;for the paper &quot; <em>Conjugate Heat Transfer Modeling of a Cold Plate Design for Hybrid-Cooled Data Centers</em>&rdquo;&nbsp;in Energies Journal.</p> <p>https://www.mdpi.com/1996-1073/16/7/3088&nbsp;</p> <p>Numerical simulations are performed using the open-source CFD code OpenFoam. Features of the numerical model described in the paper can be summarized as:</p> <ul> <li>This data set contains <em>OpenFoam</em> cases for the validation of the thermal model with the experimental data in the literature (Saitoh et al. 1993) using both <em>buoyantPimpleFoam</em> and <em>chtMultiRegionFoam</em> solvers.</li> </ul> <p><em>Saitoh, T.; Sajiki, T.; Maruhara, K. Benchmark solutions to natural convection heat transfer problem around a horizontal circular cylinder. Int. J. Heat Mass Trans. 1993, 36, 1251&ndash;1259.</em></p> <ul> <li>A multi-region unstructured mesh was created using Salome software and exported as unv files. The generated unv files can be found in the corresponding directories.</li> <li><em>Allmesh</em> script imports regions from the unv files to the <em>OpenFoam</em> and runs <em>createPatch</em> file for the application of boundary conditions .</li> <li><em>Allrun</em>&nbsp;script runs transient simulation using parallel computing.</li> <li><em>postProcess</em>&nbsp;script compares case results with experimental results and generates a plot in the directory results.</li> <li>Flow inside air and water regions are considered as laminar due to the low Reynolds numbers.</li> <li>Numerical schemes used in the solutions of constitutive equations are carefully selected to obtain results consistent with the experimental data.</li> </ul> <p>A new function object is developed for the calculation of the Nusselt number on the cylinder. This library can be downloaded via the following link and sould be&nbsp; compiled before cases are run:&nbsp;&nbsp;</p> <p><a href="https://github.com/DSTECHNO/NusseltNumber">https://github.com/DSTECHNO/NusseltNumber</a></p> <p><strong>FreeConvectionBPF.tar.gz:</strong>&nbsp;<em>OpenFoam</em> files and scripts for the transient simulation of free convection case using <em>buoyantPimpleFoam</em> solver.</p> <p><strong>FreeConvectionCHT.tar.gz:</strong>&nbsp;<em>OpenFoam</em> files and scripts for the transient simulation of free convection case using <em>chtMultiRegionFoam</em> solver.</p> <p><strong>ForcedConvectionBPF.tar.gz:</strong>&nbsp;OpenFoam files and scripts for the transient simulation of forced convection case using <em>buoyantPimpleFoam</em> solver.</p> <p><strong>ForcedConvectionCHT.tar.gz: &nbsp;</strong><em>OpenFoam</em> files and scripts for the transient simulation of forced convection case using <em>chtMultiRegionFoam</em> solver.</p>

opencc-by-4.0Feb 2023View details →
zenodo44/100

Building measured data for model validation

<p>Measured indoor/outdoor temperatures, solar radiation and heating load of a 103-m2 building in Athens, Greece. Data include measurements of two weeks, one without heating delivery to the building and another with heating delivery (with fan coils).</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

ViF-GTAD: A new Automotive Data Set with Ground Truth for ADAS/AD Development, Testing and Validation

<p>A new dataset for automated driving, which is the subject matter of this paper, identifies and addresses a gap in existing similar perception data sets. While the most state-of-the-art perception data sets primarily focus on provision of various on-board sensor measurements along with the semantic information under various driving conditions, the provided information is often insufficient since the object list and position data provided include unknown and time-varying errors. The current paper and the associated data-set describes the first publicly available perception measurement data that include not only the on-board sensor information from camera, Lidar and radar with semantically classified objects, but also the high precision ground-truth position measurements enabled by the accurate RTK assisted GPS localization systems available on both the ego vehicle and the dynamic target objects. This paper provides insight on the capturing of the data, explicitly explaining the meta data structure and the content, as well as the potential application examples where it has been, and can potentially be, applied and implemented in relation to automated driving and environmental perception systems development, testing and validation.</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Validation data set on land cover changes for RapidAI4EO project

<p>This is a reference data set collected for validation of the monthly land cover maps at a 3m and at a 10m resolution produced in the WP5. The reference data set has been collected by using Geo-Wiki toolbox for visual interpretation of very high-resolution images, including Planet data and Google maps. The data set has been collected over 3 AOIs. Each reference sample site corresponds to a 30m-by-30m box and includes information about monthly land cover type over the period 2018-2020. Land cover legend is the same as in ESA WorldCover map at a 10m resolution (https://worldcover2021.esa.int/).</p> <p>Fields:</p> <p>&quot;rowid&quot; &ndash; unique row identifier;</p> <p>&quot;sampleid&quot; &ndash; unique sample site identifier in the Geo-Wiki database;</p> <p>&quot;samplegroupid&quot; &ndash; group id with values 257(Portugal), 258 (Belgium), 259(Sicily);</p> <p>&quot;x_min&quot;,&quot;x_max&quot;,&quot;y_min&quot;,&quot;y_max&quot; &ndash; bounding box coordinates of each sample site (30m x 30m), in WGS84</p> <p>&quot;X2018_1&quot;,&quot;X2018_2&quot;,&hellip;, &quot;X2020_12&quot; &ndash; dominant land cover class in each sample site in each month from January 2018 to December 2020;</p> <p>Land cover codes:</p> <p>10 &ndash; Tree cover</p> <p>20 - Shrubland</p> <p>30 - Grassland</p> <p>40 - Cropland</p> <p>50 &ndash; Urban/built-up</p> <p>60 - Bare/Sparse vegetation</p> <p>80 - Water</p> <p>90 - Wetland</p> <p>110 - Burnt</p> <p>120 &ndash; Not sure</p>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Global Crop Type Validation Data Set for ESA WorldCereal System

<p>This dataset was created by using a new IIASA tool, called &ldquo;Street Imagery validation&rdquo; (<a href="https://svweb.cloud.geo-wiki.org/">https://svweb.cloud.geo-wiki.org/</a>) where users could check street level images (e.g., Google Street Level images, Mapillary etc.) and identify the crop type where it is possible. The advantage of this tool is that there are plenty of georeferenced images with dates, going back in time. The disadvantage is that users need to check plenty of images where only few will clearly show cropland fields that are mature enough to be identified. To make the data collection more efficient, we provided our experts with preliminary maps of points in agricultural areas where street level images are available for the year 2021. Then, the experts checked those locations in an opportunistic way. The dataset is completely independent from all the existing maps and the reference datasets.</p> <p>There are 3 main data records uploaded:</p> <ol> <li>sv_croptype_poly.zip &ndash; an archive with a shapefile containing all the collected polygons with crop type information. Not all the polygons correspond to actual field boundaries.</li> <li>sv_croptype_validations.csv &ndash; a table with crop type observations with centroid coordinates in WGS84</li> <li>sv_worldcereal_validation.csv &ndash; a table with a subset of crop type observations used in validation of WorldCereal crop type maps for 2021.</li> </ol> <p>Fields:</p> <ul> <li>&quot;id&quot; &ndash; unique observation identifier;</li> <li>&quot;imgSource&quot; &ndash; source of imagery used for visual inspection;</li> <li>&quot;imgLoc&quot; &ndash; image location;</li> <li>&quot;svImgDate&quot; &ndash; image date;</li> <li>&quot;imageIdKey&quot; &ndash; image unique identifier;</li> <li>&quot;submitedAt&quot; &ndash; date of submission of crop type observation;</li> <li>&quot;cropType&quot; &nbsp;- crop type observation;</li> <li>&quot;irrType&quot; &ndash; irrigation type;</li> <li>&quot;x&quot;, &quot;y&quot; &ndash; centroids of submitted polygons in WGS84.</li> </ul>

opencc-by-4.0Apr 2023View details →
zenodo44/100

Multicenter Validated Detection of Focal Cortical Dysplasia using Deep Learning

<p>Lesional and non-lesional patches derived from 148 FCD patients is available as a HDF5 dataset (v1.0.0; doi: 10.5061/dryad.h70rxwdgm or 10.5281/zenodo.3239446). To create this dataset, for each of the 148 FCD patients, we sampled at most 1,000 (*_N1000.h5) or 1,500 (*_N1500.h5) cortical patches (or # voxels in the lesion, whichever is lower) of size 16&times;16&times;16 within the lesion.The same number of cortical patches were sampled randomly outside the lesion. The resulting lesional and non-lesional patches were concatenated, shuffled (to add another layer of randomization), and saved along with their binary labels (compressed HDF5 dataset).</p> <p>For&nbsp;axis=1, index&nbsp;0&nbsp;is T1 and&nbsp;1&nbsp;is FLAIR.</p>

openother-openAug 2021View details →
zenodo44/100

ETHNA Validation Survey on Drivers and Barriers of Responsible Research & Innovation (RRI)

<p>This entry includes the ETHNA Validation Survey on Drivers and Barriers of Responsible Research &amp; Innovation (RRI) questions in PDF and the servey results in Excel. Ths survey was conducted by Centre for Social Innovation (ZSI) as part of the WP4 of the European ETHNA System project. The survey explores potential good RRI practices and possible measures to gauge the RRI institutionalisation progress at research-performing organisations.</p> <p>The online survey was sent in October 2021 to a broad group (10.000+) of potentially relevant expert stakeholders identified through a Web of Science database search with the aim of assessing the relevance of the identified drivers, barriers and good practices of RRI institutionalisation. Altogether 888 responses were received from 69 countries with a balanced gender representation, involving the opinion of mostly senior experts (55% having more than 15 years of experience).&nbsp;After filling out general demographic and organisational information, the respondents rated the perceived relevance of the RRI incentives, barriers and good practices that were the highest ranked at the end of the second consultation phase (using a Likert scale of 1-10).&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

TCV-X21-GENEX: influence of collisions on the validation of global gyrokinetic simulations

<p>This repository contains the data that supports the findings of the study that is published in <a href="http://doi.org/10.1063/5.0144688"><em>P. Ulbl et al., Phys. Plasmas 30, </em>052507 <em>(2023)</em></a>. Three global electromagnetic gyrokinetic simulations of the <a href="https://doi.org/10.5281/zenodo.5776286">TCV-X21</a> case have been performed. A collisionless (No Coll) simulation, one with a Bhatnagar-Gross-Krook (BGK) collision operator and one with a Fokker-Planck type Lenard-Bernstein/Dougherty (LBD) collision operator.</p> <p>For getting started, please consider the README file provided in this dataset. The Jupyter notebook herein provides a low level entry on how to access the data from the netCDF file. The zip files contain sets of input parameters that are given for future reference.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Federated Learning for Distributed Intrusion Detection Systems in Public Networks - Validation Dataset

<p>This dataset has been meticulously prepared and utilized as a validation set during the evaluation phase of &quot;Meta IDS&quot; to asses the performance of various machine learning models. It is&nbsp; now made available for interested users and researchers who seek a reliable and diverse dataset for training and testing their own custom models.</p> <p>The validation dataset comprises a comprehensive collection of labeled entries, that determines whether the packet type is &quot;malicious&quot; or &quot;benign.&quot; It covers complex design patterns that are commonly encountered in real-world applications. The dataset is designed to be representative, encompassing edge and fog layers that are in contact with cloud layer, thereby enabling thorough testing and evaluation of different models. Each sample in the dataset is labeled with the corresponding ground truth, providing a reliable reference for model performance evaluation.</p> <p>&nbsp;</p> <p>To ensure convenient distribution and storage, the dataset has been broken down into three separate batches, each containing a portion of the dataset. This allows for convenient downloading and management of the dataset. The three batches are provided as individual compressed files.</p> <p>&nbsp;</p> <p>In order to extract the data, follow the following instructions:</p> <ul> <li>Download and install bzip2 (if not already installed) from the official website or your package manager.</li> <li>Place the compressed dataset file in a directory of your choice.</li> <li>Open a terminal or command prompt and navigate to the directory where the compressed dataset file is located.</li> <li>Execute the following command to uncompress the dataset: <ul> <li>bzip2 -d filename.bz2</li> </ul> </li> <li>Replace &quot;filename.bz2&quot; with the actual name of the compressed dataset file.</li> </ul> <p>Once uncompressed, you will have access to the dataset in its original format for further exploration, analysis, and model training etc. The total storage required for extraction is approximately 800 GB in total, with the first batch requiring approximately 302 GB, the second batch requiring approximately 203 GB, and the third batch requiring approximately 297 GB of data storage.</p> <p>&nbsp;</p> <p>The first batch contains 1,049,527,992 entries, where as the second batch contains&nbsp;711,043,331 entries, and for the third and last batch we have 1,029,303,062 entries. The following table provides the feature names along with their explanation and example value once the dataset is extracted.</p> <p>&nbsp;</p> <table align="left"> <thead> <tr> <th scope="col">Feature</th> <th scope="col">Description</th> <th scope="col">Example Value</th> </tr> </thead> <tbody> <tr> <td>ip.src</td> <td>Source IP address in the packet</td> <td>a05d4ecc38da01406c9635ec694917e969622160e728495e3169f62822444e17</td> </tr> <tr> <td>ip.dst</td> <td>Destination IP address in the packet</td> <td>a52db0d87623d8a25d0db324d74f0900deb5ca4ec8ad9f346114db134e040ec5</td> </tr> <tr> <td>frame.time_epoch</td> <td>Epoch time of the frame</td> <td>1676165569.930869</td> </tr> <tr> <td>arp.hw.type</td> <td>Hardware type</td> <td>1</td> </tr> <tr> <td>arp.hw.size</td> <td>Hardware size</td> <td>6</td> </tr> <tr> <td>arp.proto.size</td> <td>Protocol size</td> <td>4</td> </tr> <tr> <td>arp.opcode</td> <td>Opcode</td> <td>2</td> </tr> <tr> <td>data.len</td> <td>Length</td> <td>2713</td> </tr> <tr> <td>eth.dst.lg</td> <td>Destination LG bit</td> <td>1</td> </tr> <tr> <td>eth.dst.ig</td> <td>Destination IG bit</td> <td>1</td> </tr> <tr> <td>eth.src.lg</td> <td>Source LG bit</td> <td>1</td> </tr> <tr> <td>eth.src.ig</td> <td>Source IG bit</td> <td>1</td> </tr> <tr> <td>frame.offset_shift</td> <td>Time shift for this packet</td> <td>0</td> </tr> <tr> <td>frame.len</td> <td>frame length on the wire</td> <td>1208</td> </tr> <tr> <td>frame.cap_len</td> <td>Frame length stored into the capture file</td> <td>215</td> </tr> <tr> <td>frame.marked</td> <td>Frame is marked</td> <td>0</td> </tr> <tr> <td>frame.ignored</td> <td>Frame is ignored</td> <td>0</td> </tr> <tr> <td>frame.encap_type</td> <td>Encapsulation type</td> <td>1</td> </tr> <tr> <td>gre</td> <td>Generic Routing Encapsulation</td> <td>&#39;Generic Routing<br> Encapsulation (IP)&rsquo;</td> </tr> <tr> <td>ip.version</td> <td>Version</td> <td>6</td> </tr> <tr> <td>ip.hdr_len</td> <td>Header length</td> <td>24</td> </tr> <tr> <td>ip.dsfield.dscp</td> <td>Differentiated Services<br> Codepoint</td> <td>56</td> </tr> <tr> <td>ip.dsfield.ecn</td> <td>Explicit Congestion<br> Notification</td> <td>2</td> </tr> <tr> <td>ip.len</td> <td>Total length</td> <td>614</td> </tr> <tr> <td>ip.flags.rb</td> <td>Reserved bit</td> <td>0</td> </tr> <tr> <td>ip.flags.df</td> <td>Don&#39;t fragment</td> <td>1</td> </tr> <tr> <td>ip.flags.mf</td> <td>More fragments</td> <td>0</td> </tr> <tr> <td>ip.frag_offset</td> <td>Fragment offset</td> <td>0</td> </tr> <tr> <td>ip.ttl</td> <td>Time to live</td> <td>31</td> </tr> <tr> <td>ip.proto</td> <td>Protocol</td> <td>47</td> </tr> <tr> <td>ip.checksum.status</td> <td>Header checksum status</td> <td>2</td> </tr> <tr> <td>tcp.srcport</td> <td>TCP source port</td> <td>53425</td> </tr> <tr> <td>tcp.flags</td> <td>Flags</td> <td>0x00000098</td> </tr> <tr> <td>tcp.flags.ns</td> <td>Nonce</td> <td>0</td> </tr> <tr> <td>tcp.flags.cwr</td> <td>Congestion Window Reduced<br> (CWR)</td> <td>1</td> </tr> <tr> <td>udp.srcport</td> <td>UDP source port</td> <td>64413</td> </tr> <tr> <td>udp.dstport</td> <td>UDP destination port</td> <td>54087</td> </tr> <tr> <td>udp.stream</td> <td>Stream index</td> <td>1345</td> </tr> <tr> <td>udp.length</td> <td>Length</td> <td>225</td> </tr> <tr> <td>udp.checksum.status</td> <td>Checksum status</td> <td>3</td> </tr> <tr> <td>packet_type</td> <td>Type of the packet which is either &quot;benign&quot; or &quot;malicious&quot;</td> <td>0</td> </tr> </tbody> </table> <p>Furthermore, in compliance with the GDPR and to ensure the privacy of individuals, all IP addresses present in the dataset have been anonymized through hashing. This anonymization process helps protect the identity of individuals while preserving the integrity and utility of the dataset for research and model development purposes.</p> <p>&nbsp;</p> <p>Please note that while the dataset provides valuable insights and a solid foundation for machine learning tasks, it is not a substitute for extensive real-world data collection. However, it serves as a valuable resource for researchers, practitioners, and enthusiasts in the machine learning community, offering a compliant and anonymized dataset for developing and validating custom models in a specific problem domain.</p> <p>&nbsp;</p> <p>By leveraging the validation dataset for machine learning model evaluation and custom model training, users can accelerate their research and development efforts, building upon the knowledge gained from my thesis while contributing to the advancement of the field.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Catalytic Rules and Validation Results for "EzMechanism: An Automated Tool to Propose Catalytic Mechanisms of Enzyme Reactions"

<p>Dataset containing the &quot;Rules of Enzyme Catalysis&quot; as created during the development of EzMechanism and the validation results of the software. For more information see https://www.biorxiv.org/content/10.1101/2022.09.05.506575v1, and the M-CSA website in https://www.ebi.ac.uk/thornton-srv/m-csa/</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Arctic shoreline displacement and validation data for two pilot study areas

<p>Arctic shoreline displacement data for two pilot study areas are&nbsp;supporting information for the paper <em>Nyl&eacute;n, Calle-Navarro and Gonzales-Inca:&nbsp;Arctic shoreline displacement with open satellite imagery and data fusion &ndash; Pilot study 1984&ndash;2022</em><em>. </em>The two study areas are:</p> <ul> <li>Tanafjorden: a&nbsp;low-arctic meso-tidal fjord coast in mainland Norway</li> <li>Ny-&Aring;lesund: a high-arctic micro-tidal glaciated coast in north-western Svalbard</li> </ul> <p>The study areas are 2500 km&sup2; each.</p> <p>The dataset includes following files:</p> <ul> <li><em>Calculating_coastal_landcover_timeseries_summaries.R</em>: code for summarizing the coastal land cover time-series in the R software.</li> <li><em>Calculating_shoreline_timeseries.R</em>: code for calculating the shoreline by fitting and smoothing a polyline to the land cover raster in the R software.</li> <li><em>X_timeseries.tif</em>: a multiband GeoTIFF raster file, with each band describing coastal land cover during one of the eight time-steps (1984&ndash;1988,&nbsp;1989&ndash;1993,&nbsp;1994&ndash;1998,&nbsp;1999&ndash;2003,&nbsp;2004&ndash;2008,&nbsp;2009&ndash;2013,&nbsp;2014&ndash;2018 and&nbsp;2019&ndash;2022).</li> <li><em>X_summary.tif</em>: a multiband GeoTIFF raster file, with bands that summarize the time-series from different viewpoints. These summary variables are: probability of belonging to the land class, long-term trend (between 1984-2003 and 2004-2022), change intensity, first time-step in water class, last time-step in water class, first time-step in land class and last time-step in land class.</li> <li><em>X_shoreline.geojson</em>: a GeoJSON vector file, consisting of polylines for the shoreline during each time-step. The attributes of the polyline layer describe the time-step and the total length of the shoreline.</li> <li><em>NyAlesund_timeseries_REDUCED.tif</em>: a reduced time-series for the Ny-&Aring;lesund study area, including only the time-steps with adequate number of observations (i.e., excluding 1984&ndash;1988, 1994&ndash;1998 and 2004&ndash;2008).</li> <li><em>X_reference_shoreline.geojson</em>: a GeoJSON vector file including the manually digitized (scale 1/5000)&nbsp;reference shoreline corresponding to the time-step&nbsp;2019&ndash;2022.</li> <li><em>X_validationpoints.geojson</em>:&nbsp;a GeoJSON vector file including&nbsp;2000 random points (within 2 km from the reference shoreline) that have been&nbsp;manually classified into water and land. The classification corresponds to the time-step&nbsp;2019&ndash;2022.</li> </ul>

opencc-by-4.0May 2023View details →
zenodo44/100

TBValid collection: Pulmonary tuberculosis validation collection

<p>The TBValid dataset comprises 870 digital patients with different profiles, each with fixed Age, BMI and MtbSputum. Each is identified by a vector of features involving biological and pathophysiological parameters to roughly represent different profiles in the population and initial bacterial load. Individual patient data collected during a clinical trial have been transformed into aggregated data, which are already irreversibly anonymised. Subsequently, these aggregated data have been sampled through the procedure described in&nbsp;&ldquo;Generation of digital patients for the simulation of tuberculosis with UISS-TB&rdquo;, doi: 10.1186/s12859-020-03776-z.&nbsp;The obtained derivative dataset, owned by its creators, does not constitute sensitive data according to European laws, and it is impossible with this dataset to re-establish the identity of the patients enrolled in the original clinical trial.</p>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Temporal Validity Change Prediction - Dataset

<p>This dataset contains data for <em>temporal validity change prediction</em>, an NLP task that will be defined in an upcoming publication. The dataset consists of five columns.&nbsp;</p> <ul> <li>target - A Tweet ID. This column must be manually rehydrated via the Twitter API to obtain the tweet text.</li> <li>follow_up - A synthetic follow-up tweet that semantically relates to the target tweet.</li> <li>context_only_tv - The expected temporal validity duration of the <strong>target&nbsp;</strong>tweet, when read in isolation.</li> <li>combined_tv - The expected temporal validity duration of the&nbsp;<strong>target&nbsp;</strong>tweet, when read <strong>together&nbsp;with the follow-up tweet</strong>.</li> <li>change - The TVCP task label, i.e., whether the temporal validity duration of the target tweet is <em>decreased</em>, unchanged&nbsp;(<em>neutral</em>), or <em>increased </em>by the information in the follow-up tweet.</li> </ul> <p>The duration labels (context_only_tv, combined_tv) are class indices of the following class distribution:<br> [no time-sensitive information, less than one minute, 1-5 minutes, 5-15 minutes, 15-45 minutes, 45 minutes - 2 hours, 2-6 hours, more than 6 hours, 1-3 days, 3-7 days, 1-4 weeks, more than one month]</p> <p>Different dataset splits are provided.</p> <ul> <li>&quot;dataset.csv&quot; contains the full dataset.</li> <li>&quot;train.csv&quot;, &quot;val.csv&quot;, &quot;test.csv&quot; contain an 80-10-10 train-val-test split.</li> <li>&quot;train[0-4].csv&quot; and &quot;test[0-4].csv&quot; respectively contain training and test data for one of 5 folds for 5-fold cross-validation. The train file contains 80% of the data, while the test file contains 20%. To replicate the original experiments, the train file should be sorted by the preprocessed target tweet text, then the first 12.5% of target tweets should be sampled to generate validation data, leading to a 70-10-20 train-val-test split.&nbsp;</li> </ul>

opencc-by-4.0Sep 2023View details →
zenodo44/100

In-situ Heating-Stage EBSD Validation of Algorithms for Prior-Austenite Grain Reconstruction in Steel

<p>High temperature EBSD and dilatometry data from the manuscript &quot;In-situ Heating-Stage EBSD Validation of Algorithms for Prior-Austenite Grain Reconstruction in Steel&quot;. This includes Gifs of the martensitic and bainitic phase transformations, individual frames as Tiff files&nbsp;and as CTF files. It also includes&nbsp;thermocouple read outs from the in-situ crucible and the raw&nbsp;data from the dilatometry experiments.</p>

opencc-by-4.0Sep 2023View details →
zenodo44/100

Functional organization of the mouse lemur Primate Microcebus murinus : from multilevel validation to comparison with humans

<p>Brain network organization in the mouse lemur (Microcebus murinus) Primate.<br> (comparison with humans)<br> Archives contain:<br> <br> - Dictionary learning analysis in mouse lemurs and humans showing networks identified in these two species.<br> <br> - Cerebral templates from mouse lemurs and humans (MNI template). They can be used to localize networks.<br> <br> - A functional atlas of the mouse lemur brain issued from resting fMRI. Resting-state functional MR images were recorded from 14 mouse lemurs at 11.7 Tesla (2 time point per animal).<br> - An atlas from human brain (issued from&nbsp;<a href="http://www.gin.cnrs.fr/fr/outils/aal-aal2/">http://www.gin.cnrs.fr/fr/outils/aal-aal...</a>) that can be used to attribute human cerebral networks.<br> <br> - Templates, atlases and networks can be easily observed together using ITK-SNAP (<a href="http://www.itksnap.org/">http://www.itksnap.org/</a>).</p> <p>if used for publication please cite:&nbsp;</p> <p><strong>Resting state functional atlas and cerebral networks in mouse lemur primates at 11.7 Tesla</strong><br> <strong>Cl&eacute;ment M Garin</strong>, Nachiket A Nadkarni, Brigitte Landeau, Ga&euml;l Ch&eacute;telat, Jean-Luc Picq, Salma Bougacha, Marc Dhenain<br> Feb 2021<br> <strong>NeuroImage</strong> 226, 117589<br> DOI: 10.1016/J.NEUROIMAGE.2020.117589<br> <a href="https://www.sciencedirect.com/science/article/pii/S1053811920310740">https://www.sciencedirect.com/science/article/pii/S1053811920310740</a></p>

opencc-by-4.0Oct 2023View details →
zenodo44/100

RnR-Exm Validation Dataset

<p>This dataset was released as part of the 2023 ISBI challenge, <a href="https://rnr-exm.grand-challenge.org/rnr-exm/">RnR-ExM</a>.&nbsp;The organizers thank Ruihan Zhang (MIT), Margaret Elizabeth Schroeder (MIT) and Chi Zhang (MIT) for contributing data to this competition.</p>

opencc-by-4.0Jan 2023View details →
zenodo40/100

Data related to the manuscript "Bayesian Calibration and Validation of a Large-scale and Time-demanding Sediment Transport Model"

<p>1) Riverbed_Elevation_Measurements.txt<br> &nbsp;&nbsp;&nbsp; Description: Measured riverbed geometry of available years<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevation 2002 [m asl], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation&nbsp;&nbsp;&nbsp;<br> &nbsp;&nbsp;&nbsp; 2013 [m asl]<br> ----------------------------------------------------------------------------------------------------------------------------<br> 2) Hydro_FT_2D_manual.txt<br> &nbsp;&nbsp; &nbsp;Description: Simulation results of the manually calibrated full model<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]</p> <p>3.1) Hydro_FT_2D_CollocationPointBase.txt<br> &nbsp;&nbsp; &nbsp;Description: Parameter combinations of the collocation point base for each of the 20 simulations conducted with the full model to&nbsp;<br> &nbsp;&nbsp;&nbsp; construct the surrogate<br> &nbsp;&nbsp; &nbsp;Rows: Critical Shields parameter, Grain Roughness, Grain Size distribution</p> <p>3.2) Hydro_FT_2D_CollocationResults.txt<br> &nbsp;&nbsp; &nbsp;Description: Simulation results of the 20 simulations conducted with the full model at the collocation points<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of simulation 1 through 20, Node ID, Easting [m asl], Northig<br> &nbsp;&nbsp;&nbsp; [m asl], Elevations 2010 [m asl] of simulation 1 through 20, Node ID, Easting [m asl], Northig [m asl], Elevations 2013 [m asl] of<br> &nbsp;&nbsp;&nbsp; simulation 1 through 20<br> ----------------------------------------------------------------------------------------------------------------------------<br> 4.1) aPC_MC_N_Combinations_Weights_prior.txt<br> &nbsp;&nbsp; &nbsp;Description: ID of prior MC runs with tested parameter combinations and corresponding importance weights<br> &nbsp;&nbsp; &nbsp;Rows: ID of MC runs, Critical Shields parameter, Grain Roughness, Grain Size distribution, importance weights<br> 4.2) aPC_MC_2005_prior.txt<br> &nbsp;&nbsp; &nbsp;Description: aPC surrogate results of prior MC runs for 2005<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of MC run 1 through 100,000<br> 4.3) aPC_MC_2010_prior.txt<br> &nbsp;&nbsp; &nbsp;Description: aPC surrogate results of prior MC runs for 2010<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevations 2010 [m asl] of MC run 1 through 100,000<br> 4.4) aPC_MC_2013_prior.txt<br> &nbsp;&nbsp; &nbsp;Description: aPC surrogate results of prior MC runs for 2013<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevations 2013 [m asl] of MC run 1 through 100,000<br> &nbsp;&nbsp; &nbsp;<br> 4.5) aPC_MC_N_Combinations_Weights_posterior.txt<br> &nbsp;&nbsp; &nbsp;Description: ID of accepted (posterior) MC runs with tested parameter combinations and corresponding importance weights<br> &nbsp;&nbsp; &nbsp;Rows: ID of accepted MC runs, Critical Shields parameter, Grain Roughness, Grain Size distribution, importance weights<br> 4.6) aPC_MC_2005_posterior.txt<br> &nbsp;&nbsp; &nbsp;Description: aPC surrogate results of posterior MC runs for 2005<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevations 2005 [m asl] of accepted MC run 1 through 857<br> 4.7) aPC_MC_2010_posterior.txt<br> &nbsp;&nbsp; &nbsp;Description: aPC surrogate results of posterior MC runs for 2010<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevations 2010 [m asl] of accepted MC run 1 through 857<br> 4.8) aPC_MC_2013_posterior.txt<br> &nbsp;&nbsp; &nbsp;Description: aPC surrogate results of posterior MC runs for 2013<br> &nbsp;&nbsp; &nbsp;Columns: Node ID, Easting [m], Northig [m], Elevation 2013 [m asl] of accepted MC run 1 through 857<br> ----------------------------------------------------------------------------------------------------------------------------<br> 5) aPC_MAP.txt<br> &nbsp;&nbsp;&nbsp; Description: Simulation results conducted with the stochastically calibrated aPC surrogate model using the MAP parameter&nbsp;<br> &nbsp;&nbsp;&nbsp; combination<br> &nbsp;&nbsp;&nbsp; Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]</p> <p>6) Hydro_FT_2D_MAP.txt<br> &nbsp;&nbsp;&nbsp; Description: Simulation results conducted with the stochastically calibrated full model using the MAP parameter combination<br> &nbsp;&nbsp;&nbsp; Columns: Node ID, Easting [m], Northig [m], Elevation 2005 [m asl], Elevation 2010 [m asl], Elevation 2013 [m asl]<br> ----------------------------------------------------------------------------------------------------------------------------<br> 7) dz.txt<br> &nbsp;&nbsp;&nbsp; Description: Riverbed Evolution for all nodes in the section of interest (n=1138) obtained with differently calibrated models for all&nbsp;&nbsp;&nbsp;&nbsp;<br> &nbsp;&nbsp;&nbsp; considered time periods<br> &nbsp;&nbsp;&nbsp; Columns: Node ID, Easting [m asl], Northig [m asl], aPC_prior 2005 [m], aPC_posterior 2005 [m], aPC_MAP 2005 [m],&nbsp;<br> &nbsp;&nbsp;&nbsp; Hydro_FT-2D_MAP 2005 [m], Hydro_FT-2D_manual 2005 [m], aPC_prior 2010 [m], aPC_posterior 2010 [m], aPC_MAP 2010 [m],<br> &nbsp;&nbsp;&nbsp; Hydro_FT-2D_MAP 2010 [m], Hydro_FT-2D_manual 2010 [m], aPC_prior 2013 [m], aPC_posterior 2013 [m], aPC_MAP 2013 [m],<br> &nbsp;&nbsp;&nbsp; Hydro_FT-2D_MAP 2013 [m], Hydro_FT-2D_manual 2013 [m]</p> <p>8) dz_CalibrationNodes.txt<br> &nbsp;&nbsp;&nbsp; Description: Riverbed Evolution for calibration nodes (n=204) obtained with differently calibrated models for all considered time<br> &nbsp;&nbsp;&nbsp; periods<br> &nbsp;&nbsp;&nbsp; Columns: Node ID, Easting [m asl], Northig [m asl], aPC_prior 2005 [m], aPC_posterior 2005 [m], aPC_MAP 2005 [m],&nbsp;<br> &nbsp;&nbsp;&nbsp; Hydro_FT-2D_MAP 2005 [m], Hydro_FT-2D_manual 2005 [m], aPC_prior 2010 [m], aPC_posterior 2010 [m], aPC_MAP 2010 [m],<br> &nbsp;&nbsp;&nbsp; Hydro_FT-2D_MAP 2010 [m], Hydro_FT-2D_manual 2010 [m], aPC_prior 2013 [m], aPC_posterior 2013 [m], aPC_MAP 2013 [m],<br> &nbsp;&nbsp;&nbsp; Hydro_FT-2D_MAP 2013 [m], Hydro_FT-2D_manual 2013 [m]</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Central Baltic EwE validation

<p>Dataset contains parameters for the Central Baltic EwE foodweb model along with forcing and validation data as well as model output. All metadata information is contained in&nbsp;ICES WGSAM REPORT 2016 (Annex 3). The report&nbsp;is uploaded with the data set or can be accessed at&nbsp;<a href="https://www.ices.dk/community/groups/Pages/WGSAM.aspx">https://www.ices.dk/community/groups/Pages/WGSAM.aspx</a></p>

opencc-by-4.0Dec 2019View details →
zenodo40/100

Model outputs for validation and inference of high‐resolution information (downscaling) of ENETwild abundance model for wild boar, January 2020 update

<p>These maps are models obtained in intermediate phases of the ENETWILD project based on available information. There are frequent updates in order to improve the results.</p> <p>Objectives:</p> <p>- Validation of previously produced hunting yield maps and new ones<br> - Downscaling to 10x10 km grid &gt;&gt;&gt; file&nbsp; &quot;January_2020_HY_nut01_10x10.tif&quot;<br> - Downscaling to 2x2 km grid &nbsp; &gt;&gt;&gt; file &quot;January_2020_HY_nut00_2x2.tif&quot;</p> <p><br> Model settings and predictors:&nbsp; &nbsp;&nbsp;<br> - Assuming cells as municipality in 10x10 km grid downscaling<br> - Assuming cells as hunting grounds in 2x2 km grid downscaling&nbsp;&nbsp; &nbsp;</p> <p>Conclusions guiding future methodological steps:<br> - To update wild boar hunting yield data for some specific regions<br> - To increase hunting yield data resolution<br> - To explore model independent parametrization for each bioregion</p> <p>For further details and methodological approach see the paper:</p> <p>ENETWILD-consortium, P. Acevedo, S .Croft, G C Smith, J. A. Blanco-Aguiar, J. Fernandez-Lopez, M. Scandura, M. Apollonio, E.Ferroglio, Oliver Keuling, M. Sange, S. Zanet, F. Brivio, T. Podg&oacute;rski, K.Petrović, G. Body, A.&nbsp; Cohen, R. Soriguer, J. Vicente (2020) Validation and inference of high-resolution information (downscaling) of ENETwild abundance model for wild boar. EFSA supporting publication 2020:EN-1787. 23pp. doi:10.2903/sp.efsa.2020.EN-1787.</p> <p>Permission for reuse hunting yield outputs is&nbsp;granted under the terms indicated&nbsp; by&nbsp;EFSA.<br> &nbsp;</p>

opencc-by-4.0Mar 2020View details →
zenodo40/100

Original dataset for "A validation of co-authorship credit models with empirical data from the contributions of PhD candidates"

<p><strong>Publication reference:</strong><br> Donner, P. (2020). A validation of co-authorship credit models with empirical data from the contributions of PhD candidates. Quantitative Science Studies, v. 1, i. 2, p. 551-564. <a href="https://doi.org/10.1162/qss_a_00048">https://doi.org/10.1162/qss_a_00048</a>.</p> <p>&nbsp;</p> <p>The file contains one row per authorship contribution statement. Rows of publications and theses are grouped.</p> <p><strong>Description of columns:</strong></p> <p>dissertation_id - an integer identifying each dissertation thesis</p> <p>university - university at which the dissertation thesis was written and PhD degree conferred</p> <p>year - publication year of the dissertation thesis</p> <p>author - dissertation thesis author name</p> <p>title - dissertation thesis title</p> <p>subject - the field of research</p> <p>publication_id - an integer identifying each publication; publication associated with more than one thesis have the same id across theses</p> <p>reference - bibliographic reference for the publication associated with the thesis</p> <p>author_count - number of authors of the publication</p> <p>author_position - position in the author byline of the credited author</p> <p>credit - claimed credit of the author in percent</p> <p>corresponding_author - flag for whether the publication author of this row is a orresponding author</p>

opencc-by-4.0Apr 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record