Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
8,038
datasets available to search
ShareScore release 0.7.1
Dataset results
8,038 results for “validation”
Diversity in the Expressed Genomic Host Response to Myocardial Infarction - Validation Dataset
<p>External validation was performed by separately hierarchically clustering 934 patients with STEMI in an independent cohort[1] into 2 groups (232 and 702 individuals) based on Illumina HT12v4-profiled PBMC expression (median time 21 hour between cardiac catheterization and blood sampling). Probes with most variable expression intensities (SD≥0.5, 216 probes, excluding ribosomal genes) were used. From the 20 most differentially expressed genes in the discovery cohort described in the manuscript Toma et al. 2022 [2], 19 were available in the validation cohort.</p> <p>Column names include the Illumina identifyer and the mapped gene name as used in the discovery cohort. Values are log2-transformed, quantile-normalized, batch-corrected values, see also [1] for methodological details.</p> <p>Acknowledgement:</p> <p>This work is supported by LIFE – Leipzig Research Center for Civilization Diseases, Universität Leipzig. LIFE is funded by means of the European Union, by the European Regional Development Fund (ERDF) and by means of the Free State of Saxony within the framework of the excellence initiative.</p> <p> </p> <p>1) Teren A, Kirsten H, Beutner F, Scholz M, Holdt LM, Teupser D, Gutberlet M, Thiery J, Schuler G, Eitel I. Alteration of multiple leukocyte gene expression networks is linked with magnetic resonance markers of prognosis after acute st-elevation myocardial infarction. <em>Scientific Reports</em>. 2017;7:41705</p> <p>2) Toma A, dos Santos C, Burzyńska B, Góra M, Kiliszek M, Stickle N, Kirsten H, Kosyakovsky L, Wang B, van Diepen S, Epelman S, Szekely Y, Marshall JC, Billia F, Lawler PR (2022), Diversity in the Expressed Genomic Host Response to Myocardial Infarction, submitted.</p>
MELA Dataset: A Benchmark for Mediastinal Lesion Analysis (Validation Set and Annotation)
<p>MELA dataset is a benchmark for developing algorithms on mediastinal lesion analysis. We hope this large-scale dataset could facilitate the research and application of automatic mediastinal lesion detection and diagnosis. </p> <p>MELA dataset contains 1100 CT scans collected from patients with one or more lesions in the mediastinum. The MELA dataset is split into a subset of 770 CT scans for training, a subset of 110 CT scans for validation, and a test set of 220 CT scans for evaluation.</p> <p>This is the Validation Set and Annotation of MELA dataset, including 110 CTs and the annotations of the whole training set and validation set. Files include:</p> <ol> <li>Val.zip: 110 CTs in NII format (nii.gz).</li> <li>mela_train_val_annotations.csv: bounding box annotations in voxel coordinates for mediastinal lesions.</li> </ol> <p> `public_id: anonymous patient ID to match images and annotations.<br> `coordX, coordY, coordZ: coordinates of the center of annotated bounding box.<br> `x_length, y_length, z_length: the length of the bounding box in three dimensions.</p>
Data for the publication "Retrieving ice-nucleating particle concentration and ice multiplication factors using active remote sensing validated by in situ observations"
<p>This repository contains the data for the paper:</p> <p>Wieder, J., Ihn, N., Mignani, C., Haarig, M., Bühl, J., Seifert, P., Engelmann, R., Ramelli, F., Kanji, Z. A., Lohmann, U., and Henneberger, J.: Retrieving ice nucleating particle concentration and ice multiplication factors using active remote sensing validated by in situ observations, Atmos. Chem. Phys. Discuss. [preprint], https://doi.org/10.5194/acp-2022-67, in review, 2022.</p> <p>More information can be found in the README files.</p> <p>Note that the scripts to reproduce the figures of the publication are available on request.</p>
Control model validation dataset
<p>The dataset is associated with the LiftWEC H2020 research project deliverables "D3.3 Tool validation and extension report" and "D4.3 Open-access experimental data from 2D LiftWEC tests". The mathematical model is based on the hypothesis and equations presented in the D3.3, "Section 4. Validation of fundamental hypothesis for global model". The model has been programmed in Python, and the lift and drag coefficients were derived from the experimental data using the method of the least squares.</p> <p>The presented data shows a very good validation agreement between the developed model and experimental data in terms of tangential and radial forces generated on hydrofoils. The estimated values of lift and drag coefficients show great potential for wave energy extraction using rotating foils.</p>
Synthetic cryo electron microscopy single particle images containing biomolecular complexes with continuous conformational variability used for validating DeepHEMNMA method and validation results
<p>This archive contains a synthetic dataset used for validating DeepHEMNMA method and the validation results. DeepHEMNMA is a deep learning extension of HEMNMA approach for analyzing continuous conformational variability of biomolecular complexes in cryo electron (cryo-EM) microscopy single particle images. We provide a training set of 20,000 images and an inference set of 50,000 images. The training images were used (1) to estimate the conformational and rigid-body parameters with HEMNMA and (2) to train the neural network using the parameters previously estimated with HEMNMA (the file with the HEMNMA-estimated parameters is provided). The inference images were used to infer the parameters with the trained neural network. Also, we provide (1) the input PDB structure, its normal modes, and the conformational and rigid-body parameters used to synthesize the 20,000 training images (ground-truth parameters) and (2) the conformational and rigid-body parameters inferred from the set of 50,000 inference images.</p> <p>The DeepHEMNMA method and the method for synthesizing images have been fully described in the following article: "Hamitouche I and Jonic S (2022), DeepHEMNMA: ResNet-based hybrid analysis of continuous conformational heterogeneity in cryo-EM single particle images. Front Mol Biosci 9, 965645. <a href="https://doi.org/10.3389/fmolb.2022.965645">https://doi.org/10.3389/fmolb.2022.965645</a> (in press)". Additionally, this article describes a test of DeepHEMNMA using one experimental cryo-EM dataset (available in EMPIAR database under the accession code EMPIAR-10016). </p>
SCEC Broadband Platform Release 22.4.0 Validation Data
<p>This package includes the set of validation plots generated with the Broadband Platform Release 22.4.</p> <p>Fabio Silva, Kevin Milner, & Philip Maechling. (2022). SCECcode/bbp: Broadband Platform Release v22.4.0 (v22.4.0). Zenodo. https://doi.org/10.5281/zenodo.7062972</p> <p>The following folders include:</p> <p>2022-08-30-bbp-part-a - Broadband validation runs using ground motions from 17 historical events.</p> <p>2022-08-30-bbp-part-a-all - Broadband validation runs using ground motions from 17 historical events (same as above, but includes additional GoF plots such as maps and distance)</p> <p>2022-08-30-bbp-part-b - Broadband verification runs against NGA-West 2 GMPEs.</p> <p>2022-08-30-bbp-converge - Convergence plots for each method and event</p> <p>2022-08-30-bbp-tables - Summary tables for each method, along with aggregate results from all methods/events per distance and period range. Also includes Dreger figure 3 plots (see references below for more information)</p>
I-CAR door crashworthiness validation
<p>During an impact, a key component to protect the occupants of a vehicle is the door. In line with the novel design and manufacturing process of vehicle doors which has been carried out in AVANGARD project, several crash tests on the developed vehicle were carried out to demonstrate the effectiveness of these actions. The results of tests were satisfactory, but especially for the doors in the case of the side impact test. Considering that it is the most critical case for the door, it is considered a great success. In both tests, due to the low weigh of the structure, the vehicle is ejected, resulting in only a small deformation of the structure. In addition, owing to the doors are very rigid, there is almost no intrusion.</p>
Validation data for a Microwave Exposure Prototype for In Vitro Growth Inhibition of Plasmodium falciparum
<p>The dataset presented in this article revolves around the validation of an innovative irradiation device designed for <em>in vitro</em> malaria tests. The motivation stems from the escalating need for new malaria treatments due to the high mortality rates, the absence of an effective vaccine, and rising antimalarial drug resistance. The theoretical background for generating this data is based on prior studies indicating that microwave exposure can non-thermally kill malaria parasites by interacting with hemozoin crystals within infected erythrocytes, a process in which healthy cells remain unaffected, probably due to the absence of these crystals. This dataset aims to provide a thorough validation of the device's biological and electrical efficacy, by providing models to simulate various parameters such as the magnetic and electric field strength, and experimentally validate the biological effects on parasite viability under controlled microwave exposure. This contribution is pivotal for advancing research in electromagnetic therapy as a potential treatment for malaria, providing a foundational dataset for future experimental and theoretical work.</p>
Trypanosoma Epitope Dataset: Valid Epitopes and Randomly Generated Peptides with Biochemical Metrics and AI-Generated Scores
<p>This dataset contains information about valid linear B-Cell epitopes from the Trypanosoma genus, as well as randomly generated peptides. It includes biochemical metrics generated by the EpiBuilder-1.0 tool and scores generated by the BepiPred-3.0 software. The data was originally collected from the IEDB and UniProtKB platforms and has been processed and enhanced with these informations for researchers interested in understanding the molecular interactions between Trypanosoma protozoans and the immune system.</p>
Satellite-derived monthly Arctic winter sea ice thickness, snow depth, freeboards, ice draft, and bulk ice density (2011-2022) and validation datasets
<h1><strong>[Description]</strong></h1> <p>This dataset is curated for a manuscript published in Earth and Space Science by Hoyeon Shi and his colleagues in April 2024. </p> <blockquote> <p>Shi, H., Tonboe, R., Lee, M., Dybkjær, G., Sohn, J., Singha, S., & Baordo, F. (2024). A Simple and Robust CryoSat-2 Radar Freeboard Correction Method Dedicated to TFMRA50 for the Arctic Winter Snow Depth and Sea Ice Thickness Retrieval. <em>Earth and Space Science</em>, <em>11</em>(10), e2024EA003715. https://doi.org/10.1029/2024EA003715</p> </blockquote> <p>Here, version 2 is uploaded, corresponding to the revised manuscript during the revision. The main changes compared to version 1 are:<br> 1) Update of the CryoSat-2 radar freeboard dataset (from v2p4 to v2p6)<br> 2) Update of the coefficients for the radar freeboard correction equations<br> 3) Extension of the retrieval period for the CS2IS2 method (April is now included)<br> 4) Removal of OIB data points used for the regression from the validation datasets<br> 5) Inclusion of the Fram Strait mooring dataset in the validation dataset</p> <p>It consists of three directories, each described below.</p> <h2><strong>01_retrieval_results</strong></h2> <p>This directory includes CryoSat-2-based monthly fields of Arctic sea ice thickness, snow depth, total freeboard, ice freeboards, ice draft, and bulk sea ice density for the winter months of the 2011-2022 period (January-March for alpha method and January-April for CS2IS2 method). Those variables are obtained using six combinations of two retrieval methods and three radar freeboard correction methods.</p> <p><em>Retrieval methods</em></p> <ul> <li>alpha method: A simultaneous retrieval method based on Shi et al. (2020) and Shi et al. (2023), combining CryoSat-2, AVHRR, and AMSR data</li> <li>CS2IS2 method: A simultaneous retrieval method based on Kwok and Marcus (2018) and Kwok et al. (2020), combining CryoSat-2 and ICESat-2 data</li> </ul> <p><em>Radar freeboard correction methods</em></p> <ul> <li>Wave speed correction method: Mallet et al. (2020)</li> <li>Empirical correction method: An empirical correction derived from the CS2_OIB_matchup data, using snow depth as a predictor</li> <li>Bias correction method: An empirical correction derived from the CS2_OIB_matchup data, doing bias correction</li> </ul> <p>The datasets used for generating this dataset are as follows:</p> <ul> <li>CryoSat-2 <br>- AWI CryoSat-2 sea ice thickness v2p6 (doi: <a href="https://doi.org/10.5281/zenodo.10044554" target="_blank" rel="noopener">10.5281/zenodo.10044554</a>)</li> <li>ICESat-2<br>- NSIDC ATL20 dataset (doi: <a href="https://doi.org/10.5067/ATLAS/ATL20.004" target="_blank" rel="noopener">10.5067/ATLAS/ATL20.004</a>)</li> <li>AVHRR<br>- Copernicus Marine Service's surface temperature datasets (doi: <a href="https://doi.org/10.48670/MOI-00130" target="_blank" rel="noopener">10.48670/MOI-00130</a>, doi: <a href="https://doi.org/10.48670/MOI-00123" target="_blank" rel="noopener">10.48670/MOI-00123</a>)</li> <li>AMSR<br>- JAXA AMSR-E 6.9 GHz brightness temperature (doi: <a href="https://doi.org/10.57746/EO.01gs73ayng11rpwk7n54aynyj1" target="_blank" rel="noopener">10.57746/EO.01gs73ayng11rpwk7n54aynyj1</a>)<br>- JAXA AMSR2 6.9 GHz brightness temperature (doi: <a href="https://doi.org/10.57746/EO.01gs73b1nzeh3g66jr4p04mr0j" target="_blank" rel="noopener">10.57746/EO.01gs73b1nzeh3g66jr4p04mr0j</a>)</li> <li>Auxiliary data<br>- Sea ice concentration: OSI SAF (doi: <a href="https://doi.org/10.15770/EUM_SAF_OSI_0013" target="_blank" rel="noopener">10.15770/EUM_SAF_OSI_0013</a>, doi: <a href="https://doi.org/10.15770/EUM_SAF_OSI_0014" target="_blank" rel="noopener">10.15770/EUM_SAF_OSI_0014</a>)<br>- Sea ice type: OSI SAF (doi: <a href="https://doi.org/10.15770/EUM_SAF_OSI_NRT_2006" target="_blank" rel="noopener">10.15770/EUM_SAF_OSI_NRT_2006</a>)</li> </ul> <p>The naming convention is 'RetrievalMethod_CorrectionMethod_yyyymm.bin'. The 'RetrievalMethod' is either 'alpha' or 'CS2IS2', and the 'CorrectionMethod' is either 'WaveSpeed,' 'Empirical,' or 'BiasCorrection.' The data format is a 32-bit floating point array in the shape of 6 x 448 x 304 (25 km polar stereographic grid). The first dimension indicates the variables (in the order of snow depth (0), sea ice thickness (1), ice freeboard (2), total freeboard (3), sea ice draft (4), and bulk sea ice density (5)). For example, to read the sea ice thickness of January 2020 based on the alpha method with an empirical correction, you may write this Python command:</p> <p><code>import numpy as np</code><br><code>data = np.fromfile('alpha_Empirical_202001.bin', dtype=np.float32).reshape(6,448,304)</code><br><code>hi = data[1,:,:]</code></p> <p>The unit of thickness-related variable is cm, and the unit of density is kg/m3. The 25 km polar stereographic grid information is available on the NSIDC website (doi: <a href="https://doi.org/10.5067/N6INPBT8Y104" target="_blank" rel="noopener">10.5067/N6INPBT8Y104</a>).</p> <h2><strong>02_valdiation data </strong></h2> <p>This directory includes reference data used for quality assessment of retrievals. There are three sub-directories:</p> <p>'Mooring_draft_psn25_monthly' includes sea ice draft measurements from the moorings in the Beaufort Sea (https://www2.whoi.edu/site/beaufortgyre/data/mooring-data/), Fram Strait (doi: <a href="https://doi.org/10.21334/npolar.2022.b94cb848" target="_blank" rel="noopener">10.21334/npolar.2022.b94cb848</a>), and the Laptev Sea (doi: <a href="https://doi.org/10.1594/PANGAEA.912927" target="_blank" rel="noopener">10.1594/PANGAEA.912927</a>, doi: <a href="https://doi.org/10.1594/PANGAEA.899275" target="_blank" rel="noopener">10.1594/PANGAEA.899275</a>).</p> <p>'OIB_SD_psn25_monthly' and 'OIB_TFB_psn25_monthly' include airborne snow depth and total freeboard measurements from NASA's Operation IceBridge campaign (doi: <a href="https://doi.org/10.5067/G519SHCKWQV6" target="_blank" rel="noopener">10.5067/G519SHCKWQV6</a>, doi: <a href="https://doi.org/10.5067/GRIXZ91DE0L9" target="_blank" rel="noopener">10.5067/GRIXZ91DE0L9</a>).</p> <p>Original data were processed to become monthly gridded data to make a comparison with satellite retrievals. The OIB data points used for the regression were excluded when processing the monthly gridded data. The naming convention of each file is 'Var_yyyymm.bin,' where 'Var' is the variable name (SD: snow depth, TFB: total freeboard, Di: ice draft). For example, you can use the following code to read the OIB snow depth in March 2014.</p> <p><code>import numpy as np</code><br><code>hs = np.fromfile('SD_201403.bin', dtype=np.float32).reshape(448,304)</code></p> <h2><strong>03_CS2_OIB_matchup</strong></h2> <p>This directory includes a match-up of AWI's CryoSat-2 L2P track data and OIB track data. The matching was done by resampling two high-resolution data on a coarser-resolution common grid (25 km polar stereographic grid) using a drop-in-a-bucket resampling method. The file format is CSV, and it is straightforward to understand when it is opened.</p> <h1><strong>[Abbreviations]</strong></h1> <p>AMSR: Advanced Microwave Scanning Radiometer<br>AVHRR: Advanced Very High Resolution Radiometer<br>AWI: Alfred Wegener Institute<br>CS2: CryoSat-2<br>JAXA: Japan Aerospace Exploration Agency<br>NASA: National Aeronautics and Space Administration<br>NSIDC: National Snow and Ice Data Center<br>OIB: Operation IceBridge<br>OSI SAF: Ocean and Sea Ice Satellite Application Facility</p> <p> </p>
Ground truth recordings for validation of spike sorting algorithms
<p><strong>Ground-truth recordings for validation of spike sorting algorithms</strong><br> </p> <p>This datasets is composed of simultaneous loose patch recordings of Ganglion Cells in mice retina, combined with dense extra-cellular recordings (252 channels). The details of the dataset can be found here <a href="https://elifesciences.org/articles/34518">https://elifesciences.org/articles/34518</a></p> <p><strong>Probe layout</strong></p> <p>The probe layout can be found as mea_256.prb. This is a 16x16 Multi Electrode Array with 30um spacing. Only 252 channels are extra-cellular signals, and the 4 corners are devoted to triggers/sync/juxta.</p> <p><strong>Struture of the data</strong></p> <p>In this dataset, you will find several individual recordings, at max 5min long each (but please do not hesitate to contact us if interested by longer recordings). The extra-cellular data are saved as 16bits unsigned integer, with a variable offset at the beginning of the file. The value of this offset is given, for every datafile, in the additional text file (padding value (see following for more details)). The files have already been filtered with a Butterworth filter of order 3 with a cut-off frequency at 100Hz</p> <p><strong>Structure of a given dataset</strong></p> <p>Please read carefully the following to understand how to load and perform spike sorting with the data. In every .tar.gz file, you will find:</p> <ul> <li> a jpg image, displaying a small chunk of the juxta-cellular signal (top left), with detected peaks and threshold. The extra-cellular spike triggered waveform, across all channels, for the juxta-spike times (top right). In the bottom, you can see the juxta-cellular spikes, for all the detected triggers (left), and on the right the voltage on the channel where the Spike Triggered Average of the extra-cellular waveform is peaking the most.</li> <li>a file .juxta.raw, as float32, with the juxta-cellular trace at 20kHz, no data offset</li> <li>a file .raw, as uint16, with the extra-cellular signals recorded for 256 channels at a sampling rate of 20kHZ. In fact, only 252 channels are extra-cellular signals, the 4 corners of the arrays are devoted to juxta-cellular and sync signals (see probe layout mea_256.prb)</li> <li>a file .triggers.npy containing the spike times of the juxta-cellular spikes, detected using a threshold of k.MAD. The exact value of k can vary on a per dataset basis, and is written in the .txt file (threshold)</li> <li>a .txt file describing some information for a given dataset, such as the threshold value used to detect the spikes, the channel in the raw file where the juxta-cellular signal is located, the minimal value of the peak for the STA (and on which channel it is located), and the header size to read the raw data</li> <li>a .params file, if you want to analyze the data with SpyKING CIRCUS</li> </ul> <p><strong>How to load the raw data in numpy</strong></p> <pre><code class="language-python">#Using the offset value from the txt file, we can load the data with memmap arrays data=numpy.memmap('mydata.raw', dtype='uint16', offset=offset, mode='r') data=data.reshape(len(data)//256, 256) #Then for example, to display the first second of channel 0 one_channel = data[:20000, 0].astype('float32') #If we want to center data around 0 one_channel -= 2**15 - 1 #And if we want to display data in micro volt, we must use the gain factor of 0.1042 provided in the header one_channel *= 0.1042</code></pre> <p> </p>
Construction, validation and application of nocturnal pollen transport networks in an agro-ecosystem: datasets collected using light microscopy and DNA metabarcoding
<p>This dataset contains all data required to reproduce the analyses conducted in Macgregor <em>et al. </em>(2018), using the R Notebook archived at doi: <a href="https://dx.doi.org/10.5281/zenodo.1322712">10.5281/zenodo.1322712</a>.</p> <p>Specifically, the dataset contains details of pollen transport detected on two matched samples, each containing 311 moths of 41 species, using two methods: a traditional light microscopy approach and a novel DNA metabarcoding approach. Both raw and manually-curated versions of each dataset are archived for full clarity. The dataset additionally contains all metadata required to fully interpret these data, including the RGB tables used to prepare Fig 4 in Macgregor <em>et al. </em>(2018).</p> <p>Macgregor <em>et al. </em>(2018) Construction, validation and application of nocturnal pollen transport networks in an agro-ecosystem: a comparison using light microscopy and DNA metabarcoding. <em>Ecological Entomology</em>, doi: <a href="https://dx.doi.org/10.1111/een.12674">10.1111/een.12674</a>.</p>
On the validity of foraminifera-based ENSO reconstructions
<p>Video of model run. Calculated oxygen isotope values of three species of planktonic foraminifera (<span class="math-tex">\( \delta^{18}O_c\)</span>) using the Foraminifera as modeled entities (FAME) module and Ocean Reanalysis data temperature and salinity data. Upper panel represents the Oceanic Nino Index (ONI) used to determine the ocean state for each monthly time step, lower panels the <span class="math-tex">\( \delta^{18}O_c\)</span> for each individual species of planktonic foraminifera (<em>G. ruber</em>; <em>G. sacculifer</em> and <em>N. dutertrei</em>).</p>
caseysaenger/ForamMgCa_PSM: files and scripts for revised version of manuscript "Calibration and validation of environmental controls on planktic foraminifera Mg/Ca using global core-top data".
<p>files and scripts for revised version of manuscript "Calibration and validation of environmental controls on planktic foraminifera Mg/Ca using global core-top data". Saenger, C. and M. N. Evans. Resubmitted to Paleoceanography and Paleoclimatology, May 3, 2019.</p>
Phenome-wide association studies across large population cohorts support drug target validation
<p>Summary-level data generated by Genomics plc as presented in:<br> Diogo, D. et al. Phenome-wide association studies across large population cohorts support drug target validation. Nat. Commun. 9, 4285 (2018). https://doi.org/10.1038/s41467-018-06540-3</p> <p>If you have any questions or comments regarding these files, please contact Genomics plc at <a href="mailto:research@genomicsplc.com">research@genomicsplc.com</a></p> <p>NOTES<br> -----------------------------<br> These analyses were carried out using the interim UK Biobank imputation data release. Analyses were restricted to a subset of "white-British" unrelated samples with a maximum sample size of 112,337 individuals. </p> <p>Case control phenotypes were defined based on categorical datafields as listed in the accompanying file. <br> Quantitative phenotypes were either rank-normalised before analysis, or beta/se values were standardised after analysis using the variance of the phenotype. The normalisation value is indicated in the accompanying file.<br> <br> All analyses included Age at assessment, sex, genotyping chip, and 10 principal components as covariates. </p> <p>We used plink1.9 linear/logistic regression as appropriate. For chromosome X variants males were treated as having 0 or 2 alternative alleles. </p> <p>The results are not adjusted for genomic control.</p> <p>DATA FILE CONTENT DESCRIPTION<br> -----------------------------<br> CHR - Chromosome<br> SNP - Variant rsID<br> ALT - Alternative allele (effect allele)<br> REF - Reference Allele (non-effect allele)<br> BP - Position in base pairs (b37, 1-based)<br> NMISS - Number of samples with non-missing genotypes<br> BETA - Effect size (log odds ratio or standardised effect size)<br> SE - Standard error<br> P - P-value<br> F_MISS - genotype missing rate<br> P_hwe - Hardy-weinberg p-value<br> MAF - ALT allele frequency</p>
Snow-Cloud Validation Masks for Multispectral Satellite Data.
<p>Geotiffs of manually validated snow, cloud, & clear-sky snow free pixels for 13 Landsat 8 images. These acquisitions are of mid-latitude mountainous regions that contain both snow and cloud cover. Four spectral libraries of snow and cloud are also provided. These are the snow and cloud spectra extracted from both these 13 scenes and the 13 L8 SPARCS Cloud Validation Masks that contained both snow and cloud. 1&2.) Snow and cloud top-of-atmosphere reflectance for the eight Landsat 8 OLI 30 meter optical bands, aggregated from the 26 scenes. 3&4.) The top-of-atmosphere reflectance for the eight Landsat 8 OLI 30 meter optical bands of all snow misidentified as cloud and cloud misidentified as snow by CFMASK, the cloud mask that ships in the BQA file of Landsat 8 Collection 1.</p>
Consensus models to predict oral rat acute toxicity and validation on a dataset coming from the industrial context
<p>We report predictive models of acute oral systemic toxicity representing a follow-up of our previous work in the framework of the NICEATM project. It includes the update of original models through the addition of new data and an external validation of the models using a dataset relevant for the chemical industry context. A regression model for LD50 and classification model for toxicity classes according to the Global Harmonized System categories were prepared. ISIDA descriptors were used to encode molecular structures. Machine learning algorithms included Support Vector Machine (SVM), Random Forest (RF) and Naïve Bayesian. Selected individual models were combined in consensus.</p> <p>The different datasets were compared using the Generative Topographic Mapping approach. It appeared that the NICEATM datasets were lacking some relevant chemotypes for chemical industry. The new models trained on enlarged data sets have applicability domain (AD) sufficiently large to accommodate industrial compounds. The fraction of compounds inside the models’ AD increased from 58 % (NICEATM model) to 94 % (new model). Yet, the increase of training sets only slightly improved of the models’ prediction performance: RMSE values decreased from 0.56 to 0.47 and balanced accuracies increased from 0.69 to 0.71 for NICEATM and new models, respectively.</p>
Visual and inertial data for validation of gliding models of ornithopters
<p>This dataset contains data from different gliding flights with an ornithopter in low wind conditions. For each experiment, the inertial information is provided.</p> <p>Additionally, the flights have been recorded from three different points of view to track and triangulate its position. The videos are provided and the position of the camera has been determined using a Leica Total Station system with submillimeter accuracy. A sample of the 2D track of the ornithopter is provided for each video and experiment. The tracking along the three cameras are synchronized.</p> <p> </p> <p>------Camera Pose structure------</p> <p> </p> <p>Three cameras with four points: three to measure orientation and the last one the lens position. The last two points are the measured fall.<br> Camera 1 -> top left, bottom left and top right.<br> Camera 2 -> top left, top rigth and bottom right.<br> Camera 3 -> top left, bottom left and top right.<br> Then there are 14 rows. The pattern is: Point1, Point2, Point3 and Lens Position.</p> <p>------IMU structure------</p> <p>time, quaternion w, quaternion x, quaternion y, quaternion z, accelerometer x, accelerometer y, accelerometer z, Gyroscope x, Gyroscope y, Gyroscope z, magnetometer x ,magnetometer y ,magnetometer z<br> units: time->ms, accelerometer->g, gyroscope->ยบ/s</p> <p> </p>
Development and validation of statistical shape models of the primary functional bone segments of the foot.
<p>This dataset comprises manually segmented three-dimensional point clouds (.STL) of magnetic resonance images of the primary functional segments of the foot - first metatarsal, midfoot (second-to-fifth metatarsals, cuneiforms, cuboid, and navicular), calcaneus, and talus. These data were used to create statistical shape models of the foot bones, utilising the GIAS2 toolbox (https://pypi.org/project/gias2/).</p>
AI validated plant observations from social media: Flickr images from central London 2011-2019
<p>This dataset is the result of using an AI image classifer to classify images of plants on social media. We believe this is the first AI validated dataset of biological records taken from social media. This represents the dawn of AI naturalists whose domain of exploration is not the outdoor world but the digital realm. These AI naturalists will trawl streams of data from all over the globe, identifying genuine images of species, and in so doing create valuable data sets that will further our understanding of the distribution of wildlife on our planet. </p> <p>This dataset contains 31,973 classifications of images taken in central London between May 2011 and September 2019 retrieved using the search term 'flower' on Flickr.com. Some images have very low classification confidence (7910 below 0.1), while others have very high confidence (3185 over 0.9). As expected given the spatial extent of the dataset many of the observations are of planted species in gardens and parks.</p> <p>August_et_al_2019.csv provides the data while metadata.txt contains a description of the data and its generation.</p> <p>An interactive visualisation of this data can be viewed at <a href="https://tomaugust.shinyapps.io/ai_flickr_data/">https://tomaugust.shinyapps.io/ai_flickr_data/</a></p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.