Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
278
datasets available to search
ShareScore release 0.9.0
Dataset results
278 results for “Validated dataset”
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Validation Set and Annotation)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is the Validation Set and Annotation of MELA dataset, including 110 CTs and the annotations of the whole training set and validation set. Files include:</p> <ol> <li>Val.zip: 110 CTs in NII format (nii.gz).</li> <li>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</li> </ol> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
Control model validation dataset
<p>The dataset is associated with the LiftWEC H2020 research project deliverables "D3.3 Tool validation and extension report" and "D4.3 Open-access experimental data from 2D LiftWEC tests". The mathematical model is based on the hypothesis and equations presented in the D3.3, "Section 4. Validation of fundamental hypothesis for global model". The model has been programmed in Python, and the lift and drag coefficients were derived from the experimental data using the method of the least squares.</p> <p>The presented data shows a very good validation agreement between the developed model and experimental data in terms of tangential and radial forces generated on hydrofoils. The estimated values of lift and drag coefficients show great potential for wave energy extraction using rotating foils.</p>
Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores
<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>
Satellite-derived monthly Arctic winter sea ice thickness, snow depth, freeboards, ice draft, and bulk ice density (2011-2022) and validation datasets
<h1><strong>[Description]</strong></h1> <p>This dataset is curated for a manuscript published in Earth and Space Science by Hoyeon Shi and his colleagues in April 2024. </p> <blockquote> <p>Shi, H., Tonboe, R., Lee, M., Dybkjær, G., Sohn, J., Singha, S., & Baordo, F. (2024). A Simple and Robust CryoSat-2 Radar Freeboard Correction Method Dedicated to TFMRA50 for the Arctic Winter Snow Depth and Sea Ice Thickness Retrieval. <em>Earth and Space Science</em>, <em>11</em>(10), e2024EA003715. https://doi.org/10.1029/2024EA003715</p> </blockquote> <p>Here, version 2 is uploaded, corresponding to the revised manuscript during the revision. The main changes compared to version 1 are:<br> 1) Update of the CryoSat-2 radar freeboard dataset (from v2p4 to v2p6)<br> 2) Update of the coefficients for the radar freeboard correction equations<br> 3) Extension of the retrieval period for the CS2IS2 method (April is now included)<br> 4) Removal of OIB data points used for the regression from the validation datasets<br> 5) Inclusion of the Fram Strait mooring dataset in the validation dataset</p> <p>It consists of three directories, each described below.</p> <h2><strong>01_retrieval_results</strong></h2> <p>This directory includes CryoSat-2-based monthly fields of Arctic sea ice thickness, snow depth, total freeboard, ice freeboards, ice draft, and bulk sea ice density for the winter months of the 2011-2022 period (January-March for alpha method and January-April for CS2IS2 method). Those variables are obtained using six combinations of two retrieval methods and three radar freeboard correction methods.</p> <p><em>Retrieval methods</em></p> <ul> <li>alpha method: A simultaneous retrieval method based on Shi et al. (2020) and Shi et al. (2023), combining CryoSat-2, AVHRR, and AMSR data</li> <li>CS2IS2 method: A simultaneous retrieval method based on Kwok and Marcus (2018) and Kwok et al. (2020), combining CryoSat-2 and ICESat-2 data</li> </ul> <p><em>Radar freeboard correction methods</em></p> <ul> <li>Wave speed correction method: Mallet et al. (2020)</li> <li>Empirical correction method: An empirical correction derived from the CS2_OIB_matchup data, using snow depth as a predictor</li> <li>Bias correction method: An empirical correction derived from the CS2_OIB_matchup data, doing bias correction</li> </ul> <p>The datasets used for generating this dataset are as follows:</p> <ul> <li>CryoSat-2 <br>- AWI CryoSat-2 sea ice thickness v2p6 (doi: <a href="https://doi.org/10.5281/zenodo.10044554" target="_blank" rel="noopener">10.5281/zenodo.10044554</a>)</li> <li>ICESat-2<br>- NSIDC ATL20 dataset (doi: <a href="https://doi.org/10.5067/ATLAS/ATL20.004" target="_blank" rel="noopener">10.5067/ATLAS/ATL20.004</a>)</li> <li>AVHRR<br>- Copernicus Marine Service's surface temperature datasets (doi: <a href="https://doi.org/10.48670/MOI-00130" target="_blank" rel="noopener">10.48670/MOI-00130</a>, doi: <a href="https://doi.org/10.48670/MOI-00123" target="_blank" rel="noopener">10.48670/MOI-00123</a>)</li> <li>AMSR<br>- JAXA AMSR-E 6.9 GHz brightness temperature (doi: <a href="https://doi.org/10.57746/EO.01gs73ayng11rpwk7n54aynyj1" target="_blank" rel="noopener">10.57746/EO.01gs73ayng11rpwk7n54aynyj1</a>)<br>- JAXA AMSR2 6.9 GHz brightness temperature (doi: <a href="https://doi.org/10.57746/EO.01gs73b1nzeh3g66jr4p04mr0j" target="_blank" rel="noopener">10.57746/EO.01gs73b1nzeh3g66jr4p04mr0j</a>)</li> <li>Auxiliary data<br>- Sea ice concentration: OSI SAF (doi: <a href="https://doi.org/10.15770/EUM_SAF_OSI_0013" target="_blank" rel="noopener">10.15770/EUM_SAF_OSI_0013</a>, doi: <a href="https://doi.org/10.15770/EUM_SAF_OSI_0014" target="_blank" rel="noopener">10.15770/EUM_SAF_OSI_0014</a>)<br>- Sea ice type: OSI SAF (doi: <a href="https://doi.org/10.15770/EUM_SAF_OSI_NRT_2006" target="_blank" rel="noopener">10.15770/EUM_SAF_OSI_NRT_2006</a>)</li> </ul> <p>The naming convention is 'RetrievalMethod_CorrectionMethod_yyyymm.bin'. The 'RetrievalMethod' is either 'alpha' or 'CS2IS2', and the 'CorrectionMethod' is either 'WaveSpeed,' 'Empirical,' or 'BiasCorrection.' The data format is a 32-bit floating point array in the shape of 6 x 448 x 304 (25 km polar stereographic grid). The first dimension indicates the variables (in the order of snow depth (0), sea ice thickness (1), ice freeboard (2), total freeboard (3), sea ice draft (4), and bulk sea ice density (5)). For example, to read the sea ice thickness of January 2020 based on the alpha method with an empirical correction, you may write this Python command:</p> <p><code>import numpy as np</code><br><code>data = np.fromfile('alpha_Empirical_202001.bin', dtype=np.float32).reshape(6,448,304)</code><br><code>hi = data[1,:,:]</code></p> <p>The unit of thickness-related variable is cm, and the unit of density is kg/m3. The 25 km polar stereographic grid information is available on the NSIDC website (doi: <a href="https://doi.org/10.5067/N6INPBT8Y104" target="_blank" rel="noopener">10.5067/N6INPBT8Y104</a>).</p> <h2><strong>02_valdiation data </strong></h2> <p>This directory includes reference data used for quality assessment of retrievals. There are three sub-directories:</p> <p>'Mooring_draft_psn25_monthly' includes sea ice draft measurements from the moorings in the Beaufort Sea (https://www2.whoi.edu/site/beaufortgyre/data/mooring-data/), Fram Strait (doi: <a href="https://doi.org/10.21334/npolar.2022.b94cb848" target="_blank" rel="noopener">10.21334/npolar.2022.b94cb848</a>), and the Laptev Sea (doi: <a href="https://doi.org/10.1594/PANGAEA.912927" target="_blank" rel="noopener">10.1594/PANGAEA.912927</a>, doi: <a href="https://doi.org/10.1594/PANGAEA.899275" target="_blank" rel="noopener">10.1594/PANGAEA.899275</a>).</p> <p>'OIB_SD_psn25_monthly' and 'OIB_TFB_psn25_monthly' include airborne snow depth and total freeboard measurements from NASA's Operation IceBridge campaign (doi: <a href="https://doi.org/10.5067/G519SHCKWQV6" target="_blank" rel="noopener">10.5067/G519SHCKWQV6</a>, doi: <a href="https://doi.org/10.5067/GRIXZ91DE0L9" target="_blank" rel="noopener">10.5067/GRIXZ91DE0L9</a>).</p> <p>Original data were processed to become monthly gridded data to make a comparison with satellite retrievals. The OIB data points used for the regression were excluded when processing the monthly gridded data. The naming convention of each file is 'Var_yyyymm.bin,' where 'Var' is the variable name (SD: snow depth, TFB: total freeboard, Di: ice draft). For example, you can use the following code to read the OIB snow depth in March 2014.</p> <p><code>import numpy as np</code><br><code>hs = np.fromfile('SD_201403.bin', dtype=np.float32).reshape(448,304)</code></p> <h2><strong>03_CS2_OIB_matchup</strong></h2> <p>This directory includes a match-up of AWI's CryoSat-2 L2P track data and OIB track data. The matching was done by resampling two high-resolution data on a coarser-resolution common grid (25 km polar stereographic grid) using a drop-in-a-bucket resampling method. The file format is CSV, and it is straightforward to understand when it is opened.</p> <h1><strong>[Abbreviations]</strong></h1> <p>AMSR: Advanced Microwave Scanning Radiometer<br>AVHRR: Advanced Very High Resolution Radiometer<br>AWI: Alfred Wegener Institute<br>CS2: CryoSat-2<br>JAXA: Japan Aerospace Exploration Agency<br>NASA: National Aeronautics and Space Administration<br>NSIDC: National Snow and Ice Data Center<br>OIB: Operation IceBridge<br>OSI SAF: Ocean and Sea Ice Satellite Application Facility</p> <p> </p>
Construction, validation and application of nocturnal pollen transport networks in an agro-ecosystem: datasets collected using light microscopy and DNA metabarcoding
<p>This dataset contains all data required to reproduce the analyses conducted in Macgregor <em>et al. </em>(2018), using the R Notebook archived at doi: <a href="https://dx.doi.org/10.5281/zenodo.1322712">10.5281/zenodo.1322712</a>.</p> <p>Specifically, the dataset contains details of pollen transport detected on two matched samples, each containing 311 moths of 41 species, using two methods: a traditional light microscopy approach and a novel DNA metabarcoding approach. Both raw and manually-curated versions of each dataset are archived for full clarity. The dataset additionally contains all metadata required to fully interpret these data, including the RGB tables used to prepare Fig 4 in Macgregor <em>et al. </em>(2018).</p> <p>Macgregor <em>et al. </em>(2018) Construction, validation and application of nocturnal pollen transport networks in an agro-ecosystem: a comparison using light microscopy and DNA metabarcoding. <em>Ecological Entomology</em>, doi: <a href="https://dx.doi.org/10.1111/een.12674">10.1111/een.12674</a>.</p>
Consensus models to predict oral rat acute toxicity and validation on a dataset coming from the industrial context
<p>We report predictive models of acute oral systemic toxicity representing a follow-up of our previous work in the framework of the NICEATM project. It includes the update of original models through the addition of new data and an external validation of the models using a dataset relevant for the chemical industry context. A regression model for LD50 and classification model for toxicity classes according to the Global Harmonized System categories were prepared. ISIDA descriptors were used to encode molecular structures. Machine learning algorithms included Support Vector Machine (SVM), Random Forest (RF) and Naïve Bayesian. Selected individual models were combined in consensus.</p> <p>The different datasets were compared using the Generative Topographic Mapping approach. It appeared that the NICEATM datasets were lacking some relevant chemotypes for chemical industry. The new models trained on enlarged data sets have applicability domain (AD) sufficiently large to accommodate industrial compounds. The fraction of compounds inside the models’ AD increased from 58 % (NICEATM model) to 94 % (new model). Yet, the increase of training sets only slightly improved of the models’ prediction performance: RMSE values decreased from 0.56 to 0.47 and balanced accuracies increased from 0.69 to 0.71 for NICEATM and new models, respectively.</p>
A Dataset of Global Land Cover Validation Samples
<p>A dataset of global land cover validation samples in 2015. In order to guarantee the confidence and objective of the validation samples, several existing reference datasets such as GLCNMO2008 training dataset, VIIRS reference dataset, STEP reference dataset, Global cropland reference data and so on, high resolution imagery in the Google earth and time-series NDVI,NDSI values of each related point are integrated to derive the global validation datasets. The dataset is provided in .shp format.</p>
QUADCOIL Validation Dataset
<p>This dataset contains the a prototype and the validation data for the stellarator coil optimization code QUADCOIL. Please extract and see <code>readme.md</code> for instructions to reproduce.</p>
Decomposition Tool Validation Dataset 1
<p>This dataset provides artifacts, benchmarks and models used to evaluate the accuracy of the performance model as well as the corresponding results:</p> <ul> <li>Deployment package of the thumbnail generation function</li> <li>50 images collected to benchmark the actual deployment</li> <li>TOSCA model with the specified open workload and performance requirement</li> <li>Predictions given by the performance model and measurements taken from the cloud platform</li> </ul>
A Synthetic Hyperspectral Dataset for Development and Validation of Phytoplankton Size Class Retrieval Models
<p><strong>A Synthetic Hyperspectral Dataset for Development and Validation of Phytoplankton Size Class Retrieval Models.</strong></p> <p>Please refer to the following scientific paper for a description of the dataset.</p> <blockquote> <p>Holtrop, T.; Van Der Woerd, H.J. (accepted) HYDROPT: An Open-Source Framework for Fast Inverse Modelling of Multi- and Hyperspectral Observations from Oceans, Coastal and Inland Waters. <em>Remote Sens. </em><strong>2021</strong>, 13, 0.</p> </blockquote>
Datasets from BUBBLES validation exercises
<p>This dataset contains the telemetry data sent by the drones during the validation exercises and the response from the BUBBLES Separation Management Environment Platform. The data was gathered during test flights, in which 14 drones performed different representative operations, including agricultural tasks, surveillance, deliveries and lifeguard operations.</p> <p>For more information about the test flights, see D5.1 Validation plan and D5.4 Validation report from BUBBLES project.</p>
Synthetic dataset used for validating MDSPACE method for analyzing continuous conformational variability of biomolecules in cryo-EM single particle images
<p>Synthetic dataset used for validating MDSPACE method for analyzing continuous conformational variability of biomolecules in cryo-EM single particle images. A README file with the contents of the dataset is included. </p>
Federated Learning for Distributed Intrusion Detection Systems in Public Networks - Validation Dataset
<p>This dataset has been meticulously prepared and utilized as a validation set during the evaluation phase of "Meta IDS" to asses the performance of various machine learning models. It is now made available for interested users and researchers who seek a reliable and diverse dataset for training and testing their own custom models.</p> <p>The validation dataset comprises a comprehensive collection of labeled entries, that determines whether the packet type is "malicious" or "benign." It covers complex design patterns that are commonly encountered in real-world applications. The dataset is designed to be representative, encompassing edge and fog layers that are in contact with cloud layer, thereby enabling thorough testing and evaluation of different models. Each sample in the dataset is labeled with the corresponding ground truth, providing a reliable reference for model performance evaluation.</p> <p> </p> <p>To ensure convenient distribution and storage, the dataset has been broken down into three separate batches, each containing a portion of the dataset. This allows for convenient downloading and management of the dataset. The three batches are provided as individual compressed files.</p> <p> </p> <p>In order to extract the data, follow the following instructions:</p> <ul> <li>Download and install bzip2 (if not already installed) from the official website or your package manager.</li> <li>Place the compressed dataset file in a directory of your choice.</li> <li>Open a terminal or command prompt and navigate to the directory where the compressed dataset file is located.</li> <li>Execute the following command to uncompress the dataset: <ul> <li>bzip2 -d filename.bz2</li> </ul> </li> <li>Replace "filename.bz2" with the actual name of the compressed dataset file.</li> </ul> <p>Once uncompressed, you will have access to the dataset in its original format for further exploration, analysis, and model training etc. The total storage required for extraction is approximately 800 GB in total, with the first batch requiring approximately 302 GB, the second batch requiring approximately 203 GB, and the third batch requiring approximately 297 GB of data storage.</p> <p> </p> <p>The first batch contains 1,049,527,992 entries, where as the second batch contains 711,043,331 entries, and for the third and last batch we have 1,029,303,062 entries. The following table provides the feature names along with their explanation and example value once the dataset is extracted.</p> <p> </p> <table align="left"> <thead> <tr> <th scope="col">Feature</th> <th scope="col">Description</th> <th scope="col">Example Value</th> </tr> </thead> <tbody> <tr> <td>ip.src</td> <td>Source IP address in the packet</td> <td>a05d4ecc38da01406c9635ec694917e969622160e728495e3169f62822444e17</td> </tr> <tr> <td>ip.dst</td> <td>Destination IP address in the packet</td> <td>a52db0d87623d8a25d0db324d74f0900deb5ca4ec8ad9f346114db134e040ec5</td> </tr> <tr> <td>frame.time_epoch</td> <td>Epoch time of the frame</td> <td>1676165569.930869</td> </tr> <tr> <td>arp.hw.type</td> <td>Hardware type</td> <td>1</td> </tr> <tr> <td>arp.hw.size</td> <td>Hardware size</td> <td>6</td> </tr> <tr> <td>arp.proto.size</td> <td>Protocol size</td> <td>4</td> </tr> <tr> <td>arp.opcode</td> <td>Opcode</td> <td>2</td> </tr> <tr> <td>data.len</td> <td>Length</td> <td>2713</td> </tr> <tr> <td>eth.dst.lg</td> <td>Destination LG bit</td> <td>1</td> </tr> <tr> <td>eth.dst.ig</td> <td>Destination IG bit</td> <td>1</td> </tr> <tr> <td>eth.src.lg</td> <td>Source LG bit</td> <td>1</td> </tr> <tr> <td>eth.src.ig</td> <td>Source IG bit</td> <td>1</td> </tr> <tr> <td>frame.offset_shift</td> <td>Time shift for this packet</td> <td>0</td> </tr> <tr> <td>frame.len</td> <td>frame length on the wire</td> <td>1208</td> </tr> <tr> <td>frame.cap_len</td> <td>Frame length stored into the capture file</td> <td>215</td> </tr> <tr> <td>frame.marked</td> <td>Frame is marked</td> <td>0</td> </tr> <tr> <td>frame.ignored</td> <td>Frame is ignored</td> <td>0</td> </tr> <tr> <td>frame.encap_type</td> <td>Encapsulation type</td> <td>1</td> </tr> <tr> <td>gre</td> <td>Generic Routing Encapsulation</td> <td>'Generic Routing<br> Encapsulation (IP)’</td> </tr> <tr> <td>ip.version</td> <td>Version</td> <td>6</td> </tr> <tr> <td>ip.hdr_len</td> <td>Header length</td> <td>24</td> </tr> <tr> <td>ip.dsfield.dscp</td> <td>Differentiated Services<br> Codepoint</td> <td>56</td> </tr> <tr> <td>ip.dsfield.ecn</td> <td>Explicit Congestion<br> Notification</td> <td>2</td> </tr> <tr> <td>ip.len</td> <td>Total length</td> <td>614</td> </tr> <tr> <td>ip.flags.rb</td> <td>Reserved bit</td> <td>0</td> </tr> <tr> <td>ip.flags.df</td> <td>Don't fragment</td> <td>1</td> </tr> <tr> <td>ip.flags.mf</td> <td>More fragments</td> <td>0</td> </tr> <tr> <td>ip.frag_offset</td> <td>Fragment offset</td> <td>0</td> </tr> <tr> <td>ip.ttl</td> <td>Time to live</td> <td>31</td> </tr> <tr> <td>ip.proto</td> <td>Protocol</td> <td>47</td> </tr> <tr> <td>ip.checksum.status</td> <td>Header checksum status</td> <td>2</td> </tr> <tr> <td>tcp.srcport</td> <td>TCP source port</td> <td>53425</td> </tr> <tr> <td>tcp.flags</td> <td>Flags</td> <td>0x00000098</td> </tr> <tr> <td>tcp.flags.ns</td> <td>Nonce</td> <td>0</td> </tr> <tr> <td>tcp.flags.cwr</td> <td>Congestion Window Reduced<br> (CWR)</td> <td>1</td> </tr> <tr> <td>udp.srcport</td> <td>UDP source port</td> <td>64413</td> </tr> <tr> <td>udp.dstport</td> <td>UDP destination port</td> <td>54087</td> </tr> <tr> <td>udp.stream</td> <td>Stream index</td> <td>1345</td> </tr> <tr> <td>udp.length</td> <td>Length</td> <td>225</td> </tr> <tr> <td>udp.checksum.status</td> <td>Checksum status</td> <td>3</td> </tr> <tr> <td>packet_type</td> <td>Type of the packet which is either "benign" or "malicious"</td> <td>0</td> </tr> </tbody> </table> <p>Furthermore, in compliance with the GDPR and to ensure the privacy of individuals, all IP addresses present in the dataset have been anonymized through hashing. This anonymization process helps protect the identity of individuals while preserving the integrity and utility of the dataset for research and model development purposes.</p> <p> </p> <p>Please note that while the dataset provides valuable insights and a solid foundation for machine learning tasks, it is not a substitute for extensive real-world data collection. However, it serves as a valuable resource for researchers, practitioners, and enthusiasts in the machine learning community, offering a compliant and anonymized dataset for developing and validating custom models in a specific problem domain.</p> <p> </p> <p>By leveraging the validation dataset for machine learning model evaluation and custom model training, users can accelerate their research and development efforts, building upon the knowledge gained from my thesis while contributing to the advancement of the field.</p>
Temporal Validity Change Prediction - Dataset
<p>This dataset contains data for <em>temporal validity change prediction</em>, an NLP task that will be defined in an upcoming publication. The dataset consists of five columns. </p> <ul> <li>target - A Tweet ID. This column must be manually rehydrated via the Twitter API to obtain the tweet text.</li> <li>follow_up - A synthetic follow-up tweet that semantically relates to the target tweet.</li> <li>context_only_tv - The expected temporal validity duration of the <strong>target </strong>tweet, when read in isolation.</li> <li>combined_tv - The expected temporal validity duration of the <strong>target </strong>tweet, when read <strong>together with the follow-up tweet</strong>.</li> <li>change - The TVCP task label, i.e., whether the temporal validity duration of the target tweet is <em>decreased</em>, unchanged (<em>neutral</em>), or <em>increased </em>by the information in the follow-up tweet.</li> </ul> <p>The duration labels (context_only_tv, combined_tv) are class indices of the following class distribution:<br> [no time-sensitive information, less than one minute, 1-5 minutes, 5-15 minutes, 15-45 minutes, 45 minutes - 2 hours, 2-6 hours, more than 6 hours, 1-3 days, 3-7 days, 1-4 weeks, more than one month]</p> <p>Different dataset splits are provided.</p> <ul> <li>"dataset.csv" contains the full dataset.</li> <li>"train.csv", "val.csv", "test.csv" contain an 80-10-10 train-val-test split.</li> <li>"train[0-4].csv" and "test[0-4].csv" respectively contain training and test data for one of 5 folds for 5-fold cross-validation. The train file contains 80% of the data, while the test file contains 20%. To replicate the original experiments, the train file should be sorted by the preprocessed target tweet text, then the first 12.5% of target tweets should be sampled to generate validation data, leading to a 70-10-20 train-val-test split. </li> </ul>
RnR-Exm Validation Dataset
<p>This dataset was released as part of the 2023 ISBI challenge, <a href="https://rnr-exm.grand-challenge.org/rnr-exm/">RnR-ExM</a>. The organizers thank Ruihan Zhang (MIT), Margaret Elizabeth Schroeder (MIT) and Chi Zhang (MIT) for contributing data to this competition.</p>
Original dataset for "A validation of co-authorship credit models with empirical data from the contributions of PhD candidates"
<p><strong>Publication reference:</strong><br> Donner, P. (2020). A validation of co-authorship credit models with empirical data from the contributions of PhD candidates. Quantitative Science Studies, v. 1, i. 2, p. 551-564. <a href="https://doi.org/10.1162/qss_a_00048">https://doi.org/10.1162/qss_a_00048</a>.</p> <p> </p> <p>The file contains one row per authorship contribution statement. Rows of publications and theses are grouped.</p> <p><strong>Description of columns:</strong></p> <p>dissertation_id - an integer identifying each dissertation thesis</p> <p>university - university at which the dissertation thesis was written and PhD degree conferred</p> <p>year - publication year of the dissertation thesis</p> <p>author - dissertation thesis author name</p> <p>title - dissertation thesis title</p> <p>subject - the field of research</p> <p>publication_id - an integer identifying each publication; publication associated with more than one thesis have the same id across theses</p> <p>reference - bibliographic reference for the publication associated with the thesis</p> <p>author_count - number of authors of the publication</p> <p>author_position - position in the author byline of the credited author</p> <p>credit - claimed credit of the author in percent</p> <p>corresponding_author - flag for whether the publication author of this row is a orresponding author</p>
Dataset for publication: Validation of large-volume batch solar reactors for the treatment of rainwater in field trials in sub-Saharan Africa, Reyneke et al. (2020). DOI: 10.1016/j.scitotenv.2020.137223
<p>Datasets, Supplementary Information and Water Safety Plan (Assessment Form and Risk Matrix) for the publication: "Validation of large-volume batch solar reactors for the treatment of rainwater in field trials in sub-Saharan Africa" which was published in Science of the Total Environment (https://doi.org/10.1016/j.scitotenv.2020.137223).</p>
Dataset for Repeated double cross validation applied to the PCA-LDA classification of SERS spectra: a case study with serum samples from hepatocellular carcinoma patients
<p>This dataset contains all the spectra used in the paper "Repeated double cross validation applied to the PCA-LDA classification of SERS spectra: a case study with serum samples from hepatocellular carcinoma patients", plus the R code to import the TXT (ASCII) files into a dataset, preprocess data, set-up and cross validate the PCA-LDA model and generate the figures shown in the paper.</p> <p>Data are available in 2 different formats: </p> <p>- 1 compressed archive ("dataset.zip") containing all the 144 TXT files (1 file = 1 spectrum) </p> <p>- 1 single CSV file (“dataset.csv”) with all the 144 spectra in the form of a table. The data are structured as follow, with each row being 1 spectrum, preceded by metadata: "acquisition_date", "substrate_batch", "class", "sample_code".</p> <p>The code for R is available as a single file "Rcode.R".</p> <p> </p>
Dataset for Millimeter-wave Mobile Sensing and Environment Mapping: Models, Algorithms and Validation
<p>Dataset of paper "Millimeter-wave Mobile Sensing and Environment Mapping: Models, Algorithms and Validation".</p> <p>The measurement data contains indoor mapping results using millimeter-wave 5G NR signals at 28 GHz. The measurement campaign was conducted in an indoor office environment in Hervanta Campus of Tampere University. Six different sets of measurements contain the range profiles after the proposed radar processing. The shared data contains the IQ data of both transmit and receive signals used during the measurement campaign.</p> <p>The file "main.m" shows how to process and plot the shared data.</p>
Datasets used in the Lab Validation of the RADON Verification Tool
<p>This repository contains the datasets that have been used to perform the lab validation of the RADON verification tool. In order to replicate the experiments, please unzip the file "validation-datasets.zip" and run the following command:</p> <pre><code class="language-bash">./run_all.rb {path_to_VT} list.json</code></pre> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.