Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
696
datasets available to search
ShareScore release 0.9.0
Dataset results
696 results for “Dataset, test,”
LLM generated Python Compiler Test Dataset
<p>This dataset is generated by integrating Large Language Models (LLMs) with AFL++ fuzzing to enhance compiler testing for CPython. It includes original Python test scripts created by LLMs such as Mistral 7B, Codellama 7B, and Gemma 7B, targeted at various compiler functionalities. These scripts were subjected to fuzzing, resulting in a rich collection of test cases that tests potential vulnerabilities. An optional minimization process with AFL-cmin refined the dataset, ensuring it focuses on test cases that significantly contribute to code coverage and bug discovery. This dataset serves as a valuable resource for improving compiler design and testing efficiency, supporting further research and development in AI-driven software testing methods.</p> <p>please see references for citations of software used in this development</p>
Dataset and code from the ATLAS phone survey to replicate analysis of HIVST positivity rates and linkage to confirmatory testing
<p>This dataset and code allow to replicate the analysis presented in the paper entitled "HIV self-testing positivity rate and linkage to confirmatory testing and care: a telephone survey in Côte d'Ivoire, Mali and Senegal" by Arsène Kra Kouassi et al. Preprint available at <a href="https://doi.org/10.1101/2023.06.10.23291206">https://doi.org/10.1101/2023.06.10.23291206</a></p><p>The data was collected within the ATLAS project, funded by Unitaid and coordinated by Solthis and IRD.</p><p>ATLAS website: <a href="https://atlas.solthis.org/">https://atlas.solthis.org/</a></p><p>ATLAS presentation on Ceped website: <a href="https://www.ceped.org/atlas">https://www.ceped.org/atlas</a></p><p>ATLAS publications portal: <a href="https://hal.science/ATLAS_ADVIH/">https://hal.science/ATLAS_ADVIH/</a></p><p> </p><p> </p>
Test dataset of VLT/SPHERE/IRDIS H2 observation Injected with simulated disks.
<p>Test data sets used in Juillard et. al. (2023) and , Juillard et. al. (2024), used to compare three different algorithms for processing data sets using ADI alone and to compare the three different strategies: RDI, ADI, and ARDI.</p> <p>This test pipeline consists of a total of 60 test data sets composed of five different disk morphologies, injected at three different contrast levels ($10^{-3}$, $10^{-4}$, and $10^{-5}$), into four different observing ADI sequences of stars without any known circumstellar signal, reflecting different observing conditions. </p> <div><strong><u>CONTENT </u></strong></div> <div> </div> <div>You will find in this folder the battery of test dataset in the compressed folder "test_cubes_sphere.zip". It contains 60 folders, one for each test dataset, with two files in each: "angle.fits" and "cube.fits"</div> <div> </div> <div>Additionaly there folders containg the empty cubes, injected disks, star flux, mask that reprensent the location of the apperture where the flux of the disk was integrated to compute contrast. The Jupyter notebook "asses_quality_of_disk_estimate.ipynb" shows how to compare a disk estimate. We also provided an example disk estimate (file 'X_3_2_0.fits') to try out the notebook.</div> <div> </div> <div>Finally, a folder named "Ref_lib_sphere" containes the different set of references frames library used in the publication Juillard et. al. (2024). In this folder, "ref_lib_X" corresponds to the optimal references for testing using empty cube X, "randref" corresponds to "shuffled" (same for every cube), and "randref_nooverlap_X" corresponds to "excluded." </div> <div>In the “Optimal” case, we used the most correlated frames, which is the same selection used in the first series, where we compare ADI, RDI, and ARDI. In the “Shuffled” case, we randomly selected a sample from the reference libraries dedicated to each of our four ADI test cubessequences into one common reference library. In this test, the random selection contains 25% of frames from the “Optimal” reference library. In the “Excluded” case, we selected for each ADI cube a sample only from references dedicated to the three other ADI cubes, creating a selection that excludes optimal references. The PCC computed for each of the three selections of references is presented in Fig. 5 of Juillard et. al. (2024).</div> <div> </div> <div><span><strong>ABOUT THE DATA</strong></span></div> <p>The data sets, obtained through the High-Contrast Data Center (HCDC), were acquired using the Infrared Dual-Band Imager and Spectrograph (IRDIS, Dohlen et al. 2008; Vigan et al. 2010) camera of the Spectro-Polarimetric High-contrast Exoplanet Research coronagraphic system on the Very Large Telescope (VLT/SPHERE, Beuzit et al. 2019). The test data sets all consist of the $H$2 channel from the dual-band $H$23 set. They were chosen to exhibit a diverse range of characteristics, including low Strehl ratio with a 26\degr\ rotation (ID #1), wind-driven halo (ID #2), an unstable speckle field (ID #3), and good Strehl ratio with an 80\degr\ field rotation (ID #4). The raw data processed with the data handling software (Pavlov et al. 2008) of the HCDC (Delorme et al. 2017), which performs dark, flat, and bad pixel correction on a coronagraphic sequence. <br>For future reference, we computed the mean and standard deviation of the Pearson correlation coefficients (PCC) between each unique pair of frames in the ADI cube. The mean PCC are as follows: Cube #1: $\mu = 0.99$; Cube #2: $\mu = 0.97$; Cube #3: $\mu = 0.93$; Cube #4: $\mu = 0.96$, with standard deviations below $0.001$ for all the cubes.</p> <p><br>The injected disks represent a range of scenarios for both debris and protoplanetary disks. As detailed in Juillard et.al 2023, this selection consists of two 75\degr\ inclined disks with varying sharpness levels (A and B), a 45\degr\ inclined disk with two concentric rings (C), a nearly face-on disk with azimuthal flux variation (D), and a hydrodynamical simulation of a disk with embedded spiral structures and a companion (E). <br>The contrast of the injected disks is determined by measuring the integrated flux within a full width at half-maximum (FWHM)-sized aperture, centered at the peak intensity of the disk, and then dividing this value by the integrated flux within an FWHM-sized aperture of the stellar point spread function. However, we made an exception for the synthetic disk E, where we measured the flux at the companion location.</p> <p>The reference frames were selected from a set of archival IRDIS observations taken with the same filter, coronagraph, and exposure time as the test data sets. These reference targets were observed between 2014 December 11 and 2021 June 1, and the raw data were calibrated through the same process as the test data sets. For each test data set, the PCC was calculated between the frames of the data set and the reference targets, excluding any observations of the data set star taken at different epochs. The PCC was calculated within a circular annulus between 0\farcs31 and 0\farcs67, which captures both the dominant speckle region and position of the waffle pattern, used for precise star centering of a coronagraphic sequence (Zurlo et al. 2014), if it was included in the observation. For each frame in the data set, the 300 best correlated reference frames were identified, and those that appeared in this selection for more than 30\% of the data set frames were selected for the final reference library. </p>
Datasets for testing the robustness of LiDAR vegetation metrics to varying point densities
<p><span>The calculation of vegetation metrics from LiDAR point clouds might be affected by the available point density of a dataset. Testing how the same LiDAR vegetation metrics differ with different point densities can therefore inform about their robustness for upscaling metrics to other areas or other LiDAR point clouds. The datasets made available here were generated to test the robustness of LiDAR vegetation metrics to varying point densities. A total of 25 LiDAR vegetation metrics representing different aspects of vegetation height, vegetation cover and structural complexity were tested (see metric definition in Kissling et al. 2023, <a href="https://doi.org/10.1016/j.dib.2022.108798">https://doi.org/10.1016/j.dib.2022.108798</a>). The metric calculation was similar to the metric calculation in the Laserchicken software (Meijer et al. 2020, <span><a href="https://doi.org/10.1016/j.softx.2020.100626">https://doi.org/10.1016/j.softx.2020.100626</a>) and the Laserfarm workflow (Kissling et al. 2022, https://doi.org/10.1016/j.ecoinf.2022.101836). The Dutch AHN4 dataset from the years 2020–2022 with a point density of 20–30 points/m<sup>2</sup> was used. A number of plots (i.e., squared polygons around centre points) were randomly placed across the Netherlands within Dutch Natura 2000 sites (using shapefiles from the European Environmental Agency). Different Dutch Natura 2000 sites were distinguished based on their dominant habitat type (dunes, grassland, marsh, shrubland, and woodland). About 100 plots were randomly placed in each habitat type. The AHN4 point cloud of each plot was clipped and then randomly downsampled to 1, 2, 5, 10, 15, 20 points per square meter, respectively. This was done for six different spatial resolutions (1, 2, 5, 10, 20 and 30 meter). The clipped points were then used to calculate the 25 LiDAR vegetation metrics for the original point density and for the six down-sampled point densities.</span></span></p>
Accompanying dataset for: A Monte Carlo Method for Metamorphic Testing of Machine Translation Services
<p>This dataset includes enhanced analysis of the machine translation data. The original dataset has been reported in [1], where white spaces were used to separate words in different languages. This is however not the best method of analyzing some Asian languages such as the Chinese language. In the present analysis, we used a character-based approach to separating the Chinese and Japanese results, hence obtaining a different set of BLEU and Cosine Similarity scores. These new scores are given in the present dataset.</p> <p>[1] Daniel Pesu, Zhi Quan Zhou, Jingfeng Zhen, & Dave Towey. (2018). Accompanying dataset for: A Monte Carlo Method for Metamorphic Testing of Machine Translation Services (Version 1.0) [Data set]. Zenodo. http://doi.org/10.5281/zenodo.1194560</p>
LimeSeg Test Datasets
<p>Image datasets from the publication : LimeSeg: A coarse-grained lipid membrane simulation for 3D image segmentation</p> <ul> <li>Vesicles.tif: spinning-disc confocal images of giant unilamellar vesicles</li> <li>HelaCell-FIBSEM.tif: a 3D Electron Microscopy (EM) dataset of nearly isotropic sections of a Hela cell, acquired with a focused ion beam scanning electron microscope (FIB-SEM). Sections are aligned with TrackEm2 (doi: ), without additional preprocessing.</li> <li>DrosophilaEggChamber.tif: point scanning confocal images of a Drosophila egg chamber. Channel 1: cell nuclei stained with DAPI. Channel 2: cell membranes visualized with fused membrane proteins Nrg::GFP and Bsg::GFP. </li> </ul> <p>Image metadata contains extra information including voxel sizes.</p> <p> </p>
Test Phenology Annotations From Herbarium Specimen Dataset - Prunus serotina
<p>This is a test dataset of phenology annotations for the Black Cherry, P. serotina, used in the submitted manuscript:</p> <p>Brenskelle, L., B. Stucky, J. Deck, R. Walls, R. P. Guralnick [submitted]. Integrating herbarium specimen observations into global phenology data systems. Applications in Plant Sciences.</p>
Cellulose test MD dataset for analysis with Galaxy and BRIDGE
<p>This coordinate and trajectory dataset is that of cellulase and octaose substrate in water. It has been derived from the <a href="https://www.rcsb.org/structure/7cel">7CEL PDB</a> structure of a fungal cellobiohydrolase. </p> <p>The original enzyme has been modified to revert the mutation at position 217 and to include disulfide bonds. The octaose substrate is an oligosaccharide consisting of 8 beta 1-4 linked glucose monomers. The enzyme and substrate have been placed in a cubic TIP3P water box containing 0.15 M ions (NaCl).</p> <p>The files includes are: </p> <ul> <li><strong>cbh1test.pdb</strong>: the coordinates of the entire system (water, protein and substrate) in PDB format. </li> <li><strong>cbh1test.dcd</strong>: a short MD trajectory (16 frames) in CHARMM DCD format.</li> </ul>
3D IQ Test Task (3D-IQTT) - A Dataset for Quantitative Evaluation of 3D Reconstruction from 2D Images
<p>3D reconstruction is mostly evaluated qualitatively. With this dataset, we are introducing a new difficult quantitative task, the 3D IQ test task (3D-IQTT).</p> <p>It is designed to be similar to mental rotation questions found in some IQ tests. Each element in the dataset consists of 4 images: reference object and answers 1-3. One of the answers is the reference object but randomly rotated. For every question, dataset users have to use their model to pick the rotated model out of the 3 possible answers.</p> <p>The dataset encourages semi-supervised or unsupervised 3D reconstruction because it contains a large corpus of unlabeled data and only a small set of labeled data where the correct answer is known.</p> <p>All the images are of blocky 3D shapes floating in space in front of a black background.</p> <p>Demo scripts for loading/processing the dataset can be found at <a href="https://github.com/fgolemo/3D-IQTT">https://github.com/fgolemo/3D-IQTT</a></p> <p>The dataset consists of:</p> <ul> <li> <pre>3diqtt-v2-train.h5 (XZ-compressed)</pre> <strong>(Training Dataset)</strong> <ul> <li> <pre>/labeled</pre> <ul> <li> <pre>/questions</pre> format: [10,000 x 4 x 128 x 128 x 3], corresponding to (10k items) x (reference + 3 answers) x (img width) x (img height) x (RGB), np.float32 in range [0,1]</li> <li> <pre>/answers</pre> format: [10,000], corresponding to (10k answers), np.uint8, one of the following three items: [0,1,2]</li> </ul> </li> <li> <pre>/unlabeled</pre> <ul> <li> <pre>/questions</pre> format: [100,000 x 4 x 128 x 128 x 3], corresponding to (100k items) x (reference + 3 answers) x (img width) x (img height) x (RGB), np.float32 in range [0,1]</li> </ul> </li> </ul> </li> <li> <pre>3diqtt-v2-test.h5</pre> <strong>(Test Dataset)</strong> <ul> <li> <pre>/questions</pre> format: [10,000 x 4 x 128 x 128 x 3], corresponding to (10k items) x (reference + 3 answers) x (img width) x (img height) x (RGB), np.float32 in range [0,1].<br> <strong>Important! This is what you have to evaluate yourself on. We have the correct answers but they are not public.</strong></li> </ul> </li> <li> <pre>3diqtt-v2-val.h5</pre> <strong>(Validation Dataset)</strong> <ul> <li> <pre>/questions</pre> format: [10,000 x 4 x 128 x 128 x 3], corresponding to (10k items) x (reference + 3 answers) x (img width) x (img height) x (RGB), np.float32 in range [0,1]</li> <li> <pre>/answers</pre> format [10,000], corresponding to (10k answers), np.uint8, one of the following three items: [0,1,2]</li> </ul> </li> </ul> <p> </p> <p><strong>Important:</strong> Before use, the main training dataset (3diqtt-v2-train.h5.xz) needs to be decompressed. This can take up to 24h depending on your hardware. We apologize for any inconvenience caused by this. The uncompressed file has a size of ~74GB. The reason for this compression was a restriction on the size of individual files. The command for decompression is "<strong>unxz</strong><strong> 3diqtt-v2-train.h5.xz</strong>" on Unix machines.</p> <p><strong>If you use this dataset, please cite it.</strong></p>
Metrics for two-sample tests: results on JetNet dataset
<p>The repository includes version 1.0 (v1.0) of the code and results corresponding to the GitHub repository <a href="https://github.com/TwoSampleTests/JetNetMetrics" target="_blank" rel="noopener">JetNetMetrics</a>.</p> <p>Publishing information and arXiv identifier will be added after publication of the main manuscript related to the data.</p>
SwinUnet_HLS_CLOUD_test_dataset
<p>Test dataset for Swin-Unet model trained with Sentinel 2 surface reflectance from Harmonized Landsat Sentinel-2 (HLS)</p>
Codes and test datasets developed for Mapping paleolacustrine deposits with a UAV-borne multispectral camera: Implications for future drone mapping on Mars.
<p>NASA’s Ingenuity Mars Helicopter has ushered in a new era in planetary exploration by utilizing Unmanned Aerial Vehicles (UAVs) to enhance our understanding of planetary surfaces. This project evaluates the potential of UAVs for mapping Martian environments, using Lake Natron, Tanzania, as an analog for Martian paleolakes.</p> <p>During two field seasons (January and July 2023), we employed a Phantom 4 Pro drone equipped with a MicaSense RedEdge-M multispectral camera and a TerraSpec Halo VNIR-SWIR spectrometer to capture high-resolution imagery and spectral data. Almost all image processing and analysis were performed using Python scripting, except for image mosaic and Digital Elevation Model (DEM) generation.</p> <p>We benchmarked the onboard image processing capabilities using a Raspberry Pi 5 single-board computer. </p> <p>In this repository, we share all the code developed during our study. Processing steps include,<br>1. DN to radiance conversion<br>2. Panel radiance extraction<br>3. Calculate reflectance factors using DLS data<br>4. Calculate reflectance at MicaSense band<br>5. Convert radiance to reflectance using 1 point empirical line method (1p ELM)<br>6. Convert radiance to reflectance using 2 point empirical line method (2p ELM)<br>7. Atmospheric correction using 6SV method<br>8. Convert radiance to reflectance using DLS data<br>9. Calculate Band indices<br>10. Weighted Kmean clustering<br>11. Finding the optimal number of clusters using the elbow method<br>12. Cmean clustering</p> <p>We also included sample image data used in the study. Feel free to contact us for more information/data.</p>
DUAL-T User Testing Dataset
<p><em><strong>The file</strong></em></p> <p>This file contains keylogging data from 21 professional literary translators, who translated three short stories in three different workflows. Data was collected between October 2023 and January 2024.</p> <p> </p> <p><em><strong>Participants</strong></em></p> <p>Participants have been assigned codes from <strong>P01 to P23</strong>. Originally, there were 24 participants, however logs for P04a, P09, and P20 had to be excluded from analysis due to technical errors, so they are not present in the spreadsheet.</p> <p> </p> <p><em><strong>Workflows</strong></em></p> <p>The workflows were:</p> <p><strong>WF01</strong>: Microsoft Word</p> <p><strong>WF02</strong>: Trados Studio 2022</p> <p><strong>WF03</strong>: Machine-Translation Postediting Platform (proprietary tool, not available on the market)</p> <p>Participants had access to an internet browser (Chrome) for all three conditions.</p> <p> </p> <p><em><strong>Source texts</strong></em></p> <p>The texts used for the experiments were taken from the short story collection <em>One More Thing</em>, by BJ Novak (2014). The three short stories used are:<br><br><strong>T01</strong>: Rome</p> <p><strong>T02</strong>: The Beautiful Girl in the Bookstore</p> <p><strong>T03</strong>: They Kept Driving Faster and They Outrun the Rain</p> <p> </p> <p><strong><em>The keylogging data</em></strong></p> <p>Keylogging data was collected using Inputlog 8.0.0.17 (<a href="https://www.inputlog.net/">https://www.inputlog.net/</a>)</p> <p>The data available in the spreadsheet includes:</p> <ul> <li>Translation time (hh:mm:ss)</li> <li>Translation time (s)</li> <li>Translation time (m)</li> <li>Total <em>n</em> keystrokes</li> <li>Total <em>n</em> mouse actions</li> <li>Total <em>n</em> pauses</li> <li>Pause duration (s)</li> <li>Mean duration of pauses (s)</li> <li>Pause ratio</li> <li>Time spent inside tool (s)</li> <li>Time spent outside tool (s)</li> <li>Source text <em>n</em> words</li> <li>Source text <em>n</em> characters</li> <li>Target text <em>n</em> words</li> <li>Target text <em>n</em> characters</li> <li>Seconds per source text character</li> </ul>
FeM dataset – An iron ore labeled images dataset for segmentation training and testing
<p>This dataset is composed of 81 pairs of correlated images. Each pair contains one image of an iron ore sample acquired through reflected light microscopy (RGB, 24-bit), and the corresponding binary reference image (8-bit), in which the pixels are labeled as belonging to one of two classes: ore (0) or embedding resin (255).</p> <p>The sample came from an itabiritic iron ore concentrate from Quadrilátero Ferrífero (Brazil) mainly composed of hematite and quartz, with little magnetite and goethite. It was classified by size and concentrated with a dense liquid. Then, the fraction -149+105 μm with density greater than 3.2 was cold mounted with epoxy resin and subsequently ground and polished.</p> <p>Correlative microscopy was employed for image acquisition. Thus, 81 fields were imaged on a reflected light microscope with a 10× (NA 0.20) objective lens and on a scanning electron microscope (SEM). In sequence, they were registered, resulting in images of 999×756 pixels with a resolution of 1.05 µm/pixel. Finally, the images from SEM were thresholded to generate the reference images.</p> <p>Further description of this sample and its imaging procedure can be found in the work by Gomes and Paciornik (2012).</p> <p>This dataset was created for developing and testing deep learning models on semantic segmentation tasks. The paper of Filippo et al. (2021) presented a variant of the DeepLabv3+ model that reached mean values of 91.43% and 93.13% for overall accuracy and F1 score, respectively, for 5 rounds of experiments (training and testing), each with a different, random initialization of network weights.</p> <p>For further questions and suggestions, please do not hesitate to contact us.</p> <p> </p> <p><strong>Contact email</strong>: ogomes@gmail.com</p> <p> </p> <p>If you use this dataset in your own work, please cite this DOI: 10.5281/zenodo.5014700</p> <p> </p> <p>Please also cite this paper, which provides additional details about the dataset:</p> <p>Michel Pedro Filippo, Otávio da Fonseca Martins Gomes, Gilson Alexandre Ostwald Pedro da Costa, Guilherme Lucio Abelha Mota. <em>Deep learning semantic segmentation of opaque and non-opaque minerals from epoxy resin in reflected light microscopy images</em>. <strong>Minerals Engineering</strong>, Volume 170, 2021, 107007, https://doi.org/10.1016/j.mineng.2021.107007.</p> <p> </p>
Deconvolution Test Dataset
<p>This a test dataset, HeLa cells stained for action using Phalloidin-488 acquired on confocal Zeiss LSM710, which contains</p> <p>- Ph488.czi (contains all raw metadata)</p> <p>- Raw_large.tif ( is the tif version of Ph488.czi, provided for conveninence as tif doesn't need Bio-Formats to be open in Fiji )</p> <p>- Raw.tif , is a crop of the large image</p> <p>- PSFHuygens_confocal_Theopsf.tif , is a theoretical PSF generated with HuygensPro</p> <p>- PSFgen_WF_WBpsf.tif , is a theoretical PSF generated with PSF generator</p> <p>- PSFgen_WFsquare_WBpsf.tif, is the result of the square operation on PSFgen_WF_WBpsf.tif , to approximate a confocal PSF</p>
LiftWEC deliverable D4.3: Dataset from 2D experimental test campaign
<p><em>This dataset contains 2-dimensional wave tank testing data for a wave-driven rotating hydrofoil model. The model tested is composed of one or two hydrofoils rotating around a horizontal axis, perpendicular to the wave direction. The model was tested in a range of regular and irregular seas. The data contains measurements of the model in the wave tank including; wave measurement, rotor position, forces on the hydrofoils, and torque on the power take off. This data is the first of two sets of wave tank data generated for the LiftWEC H2020 research project. This first set consists of results for the device tested in 2D, while the second set will contain results for tests conducted in 3D. "LiftWEC Deliverable D4.3 Report on 2D experimental testing dataset" describes this dataset and for a complete description of the test campaign, readers are directed to "LiftWEC Deliverable D4.4. </em> Report on physical modelling of 2D LiftWEC concepts <em>"</em></p>
Dataset from: A test of the reproductive assurance hypothesis in Ipomoea hederacea: does inbreeding depression counteract the benefits of self-pollination?
<p><strong>PREMISE: Darwin proposed that self-pollination in allegedly outcrossing species might act as a reproductive assurance mechanism when pollinators or mates are scarce; however, in natural populations, the benefits of selfing may be opposed by seed discounting and inbreeding depression. While empirical studies show variation among species and populations in the magnitude of reproductive assurance, little is known about the counterbalancing effects of inbreeding depression.</strong></p> <p><strong>METHODS: By comparing the female reproductive success of emasculated and open-pollinated flowers, we assessed the reproductive assurance hypothesis in two Mexican populations of <em>Ipomoea hederacea.</em> In one population we assessed temporal variation in reproductive assurance for three years. We evaluated inbreeding depression on seed production, seedling germination, and dry plant mass by contrasting self- and cross-hand pollination treatments in one population for two years.</strong></p> <p><strong> KEY RESULTS: The contribution of self-pollination to female reproductive success was high and consistent between populations, but there was variation in reproductive assurance across years. Inbreeding depression was absent in the early stages of progeny development, but there was a small negative effect of inbreeding in the probability of germination and the mass of adult progeny. </strong></p> <p><strong>CONCLUSIONS: Self-pollination provided significant reproductive assurance in <em>I. hederacea </em>but this contribution was variable across time. The contribution of reproductive assurance is probably reduced by inbreeding depression in later stages of progeny development, but this counter-effect was small in the study populations. This study supports the hypothesis that reproductive assurance with limited inbreeding depression is likely an important selective force in the evolution of self-pollination in the genus <em>Ipomoea</em>. </strong></p>
IEEE New England 39-bus test case: Dataset for the Transient Stability Assessment
<p>The <strong>dataset</strong> contains <strong>350</strong> <strong>features</strong> engineered from the phasor measurements (PMU-type) signals from the <strong>IEEE New England 39-bus power system</strong> test case network, which are generated from the 9360 systematic MATLAB®/Simulink electro-mechanical transients simulations. It was prepared to serve as a convenient and open database for experimenting with different types of <strong>machine learning</strong> techniques for <strong>transient stability assessment</strong> (TSA) of electrical power systems.</p> <p>Different load and generation levels of the New England 39-bus benchmark power system were systematically covered, as well as all three major types of short-circuit events (three-phase, two-phase and single-phase faults) in all parts of the network. The consumed power of the network was set to 80%, 90%, 100%, 110% and 120% of the basic system load levels. The short-circuits were located on the busbar or on the transmission line (TL). When they were located on a TL, it was assumed that they can occur at 20%, 40%, 60%, and 80% of the line length. Features were obtained directly from the time-domain signals at the pickup time (pre-fault value) and at the trip time (post-fault value) of the associated distance protection relays.</p> <p>This is a <strong>stochastic dataset</strong> of 3120 cases, created from the population of 9360 systematic simulations, which features a statistical distribution of different fault types, as follows: single-phase (70%), double-phase (20%) and three-phase faults (10%). It also features a <strong>class imbalance</strong>, with less than 20% of cases belonging to the unstable class. Dataset is a compressed CSV file.</p> <p><strong>List of feature names in the dataset:</strong></p> <ul> <li><em>WmGx</em> - rotor speed for each generator Gx, from G1 to G10,</li> <li><em>DThetaGx</em> - rotor angle deviation for each generator Gx, from G1 to G10,</li> <li><em>ThetaGx</em> - rotor mechanical angle for each generator Gx, from G1 to G10,</li> <li><em>VtGx</em> - stator voltage for each generator Gx, from G1 to G10,</li> <li><em>IdGx</em> - stator d-component current for each generator Gx, from G1 to G10,</li> <li><em>IqGx</em> - stator q-component current for each generator Gx, from G1 to G10,</li> <li><em>LAfvGx</em> - pre-fault power load angle for each generator Gx, from G1 to G10,</li> <li><em>LAlvGx</em> - post-fault power load angle for each generator Gx, from G1 to G10,</li> <li><em>PfvGx</em> - pre-falut value of the generator active power for each generator Gx, from G1 to G10,</li> <li><em>PlvGx</em> - post-falut value of the generator active power for each generator Gx, from G1 to G10,</li> <li><em>QfvGx</em> - pre-falut value of the generator reactive power for each generator Gx, from G1 to G10,</li> <li><em>QlvGx</em> - post-falut value of the generator reactive power for each generator Gx, from G1 to G10,</li> <li><em>VAfvBx</em> - pre-fault bus voltage magnitude in phase A for each bus Bx, from B1 to B39,</li> <li><em>VBfvBx</em> - pre-fault bus voltage magnitude in phase B for each bus Bx, from B1 to B39,</li> <li><em>VCfvBx</em> - pre-fault bus voltage magnitude in phase C for each bus Bx, from B1 to B39,</li> <li><em>VAlvBx</em> - post-fault bus voltage magnitude in phase A for each bus Bx, from B1 to B39,</li> <li><em>VBlvBx</em> - post-fault bus voltage magnitude in phase B for each bus Bx, from B1 to B39,</li> <li><em>VClvBx</em> - post-fault bus voltage magnitude in phase C for each bus Bx, from B1 to B39,</li> <li><em>Stability</em> - binary indicator (0/1) that determines if the power system was stable or unstable (0 - stable, 1 - unstable); this is the label variable.</li> </ul> <p><strong>License</strong>:<strong> </strong>Creative Commons CC-BY.</p> <p><strong>Disclaimer</strong>: This dataset is provided "as is", without any warranties of any kind.</p>
BinauRec: A dataset to test the influence of the use of room impulse responses on binaural speech enhancement
<p>BinauRec is a dataset for binaural speech enhancement. It is composed of real recordings, measured and simulated room impulse responses for the same audio scenes. Measurements are realized using behind-the-ears hearing aid shells, with and without a dummy head.</p>
ConFiRMa dataset_04: simulation of tests on CRM strengthened masonry buildings with the OOFEM code (intermediate, multi-layer level modelling)
<p>The Dataset collects the input files developed for the simulation of tests on one, two and three stroeys masonry buildings strengthened through Composite Reinforced Mortar with the free open-source code OOFEM (intermediate, multi-layer level modelling).</p> <p>OOFEM Version 2.5 (https://doi.org/10.5281/zenodo.4339630) was used for running the analyzes.</p> <p>ReadMe file provides a description of the different input files.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.