Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
252
datasets available to search
ShareScore release 0.9.0
Dataset results
252 results for “Synthetic Data”
Wind measurement data from the publication: "Development of a load model validation framework applied to synthetic turbulent wind field evaluation"
<h3>Dataset description:</h3> <p>This datasat represents supplementary material used in the contribution "Development of a load model validation framework applied to<br>synthetic turbulent wind field evaluation" by Meyer, Huhn and Gottschall.</p> <p>Wind measurements from the Testfeld BHV are made available. For installation details, see the mentioned reference.</p> <p> </p> <h3>File description:</h3> <ul> <li>Lidar_HWS.nc - Horizontal wind speed measurements (10 min averages) from a WindCube V2 vertical profiler for one day with a low-level jet occurrence ( <div> <div>2021-04-20)</div> </div> </li> <li>Cups_HWS.nc - Horizontal wind speed measurements (10 min averages) from cup anemometer installed on a met mast for the same day</li> <li>Ensemble_averaged_Spectra.nc - Ensemble averaged spectra for neutral and near neutral situations from a Gill Windmaster at 110m above ground level, used to fit the Mann and KSEC model parameters</li> </ul> <h3> </h3> <h3>Referencing:</h3> <p>When used, please cite like the following:</p> <p>Meyer, Paul J., Matthias L. Huhn, and Julia Gottschall. 2024. "Development of a Load Model Validation Framework Applied to Synthetic Turbulent Wind Field Evaluation" <em>Energies</em> 17, no. 4: 797. https://doi.org/10.3390/en17040797</p> <p> </p> <p> </p>
VoroCrack3d: An annotated data set of 3d CT concrete images with synthetic crack structures
<p>VoroCrack3d is an annotated data set of 3d CT images of concrete with synthetic crack structures. Its main purpose is the training and testing of machine learning models for 3d crack segmentation. The data set comprises 1344 images together with their corresponding ground truths. The concrete backgrounds are cropped out sections of size 400x400x400 voxels of CT images of concrete. To this end, several different concrete samples were scanned (normal concrete (NC), high-performance concrete (HPC), ultra-high-performance concrete (UHPC), air pore concrete; without and with reinforcements (straight steel fibers, crimped steel fibers, hooked-end steel fibers, polypropylene fibers, fibers made of glass fiber-reinforced polymer). The original concrete images have a resolution between 2.8 and 106 micrometers.</p> <p>The crack structures are modeled via minimum-weight surfaces in Voronoi diagrams according to the paper</p> <p>[1] C. Jung, C. Redenbach, Crack Modeling via Minimum-Weight Surfaces in 3d Voronoi Diagrams, Journal of Mathematics in Industry, 13, 10 (2023). https://doi.org/10.1186/s13362-023-00138-1.</p> <p>The surfaces are discretized, dilated and superimposed on the concrete backgrounds.</p> <p>The data set offers a high variety regarding concrete types, noise levels and crack widths, shapes, regularity and branching. This makes it suitable for studying the generalizability and robustness of 3d crack segmentation methods.</p> <p>______________________________________________________________________________________________</p> <p>The folder 'data' contains seven subfolders, each containing the data generated from a specific concrete type (NC, HPC, air pore concrete, polypropylene fiber-reinforced concrete, steel fiber-reinforced concrete (straight, crimped and hooked-end steel fibers)).</p> <p>Each subfolder again contains four subfolders according to the point process model that was used for generating the 3d Voronoi diagrams. The point processes and Voronoi diagrams are restricted to windows of size 400x150x400. </p> <p>- 'hc': Hard core point process with 60% volume density and intensity 0.000025 obtained from force-biased sphere packing.<br>- 'matclust': Matérn cluster process with parent intensity 0.0002/50, offspring intensity 50 and cluster radius 20.<br>- 'ppp': Poisson point process with intensity 0.0002.<br>- 'ppp-scaled': Poisson point process with intensity 0.0002 (but inside 200x150x200 window). The resulting Voronoi diagram is stretched in x- and z- direction by a factor of 2.</p> <p>Each of these contains five subfolders: one for the 3d input images, two for the corresponding labels (ground truths; one with and one without pores/fibers), one for the input and label previews (slice z=200 for each of the images) and a misc folder containing the concrete background without crack and, if applicable, the pore/fiber segmentation image.</p> <p>The data itself then contains 48 images:<br>1a-1d: crack with up to seven branches; fixed crack width (~1 voxel).<br>2a-2d: crack with up to four branches; fixed crack width (~1 voxel).<br>3a-3d: crack with up to one branch; fixed crack width (~1 voxel).<br>4a-4d: crack with no branches; fixed crack width (~1 voxel).<br>5a-5d: crack with no branches; fixed crack width (~3 voxels).<br>6a-6d: crack with no branches; fixed crack width (~5 voxels).<br>7a-7d: crack with no branches; fixed crack width (~7 voxels).<br>8a-8d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.01);<br>9a-9d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.02);<br>10a-10d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.05);<br>11a-11d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.1);<br>12a-12d: crack with up to seven branches; multiscale crack (bernoulli parameter 0.2);</p> <p>The names 'a'-'d' indicate level of added noise added to the image:<br>a: None.<br>b: Uniformly on [-sigma,sigma] <br>c: Uniformly on [-2*sigma,2*sigma] <br>d: Uniformly on [-4*sigma,4*sigma] <br>Negative values are mapped to 0. <br>For inputs of type int, noise values are rounded to the nearest integer.<br>(sigma = standard deviation of voxel greyvalues in image)</p> <p>Note that the grey values in the ground truths correspond to the local crack width. They can be thresholded to obtain binary masks.</p> <p>For more details, we refer to [1].</p>
Synthetic time series data generation for edge analytics
<p>In this research, we create synthetic data with features that are like data from IoT devices. We use an existing air quality dataset that includes temperature and gas sensor measurements. This real-time dataset includes component values for the Air Quality Index (AQI) and ppm concentrations for various polluting gas concentrations. We build a JavaScript Object Notation (JSON) model to capture the distribution of variables and structure of this real dataset to generate the synthetic data. Based on the synthetic dataset and original dataset, we create a comparative predictive model. Analysis of synthetic dataset predictive model shows that it can be successfully used for edge analytics purposes, replacing real-world datasets. There is no significant difference between the real-world dataset compared the synthetic dataset. The generated synthetic data requires no modification to suit the edge computing requirements. The framework can generate correct synthetic datasets based on JSON schema attributes. The accuracy, precision, and recall values for the real and synthetic datasets indicate that the logistic regression model is capable of successfully classifying data</p>
Synthetic datasets reflecting the shRNA-seq knockdown ENCODE data for HepG2 and K562 with coresponding GRN
<p>Synthetic data correspond to the ENCODE data for cell lines HepG2 (https://www.encodeproject.org/biosamples/ENCBS282XVK/) and K562 (https://www.encodeproject.org/biosamples/ENCBS023XVB/). The data and networks were generated using GeneSPIDER (publicly available at https://bitbucket.org/sonnhammergrni/genespider/).</p> <p> </p> <p><strong>Table.1 </strong>Description of the files</p> <table> <tbody> <tr> <td>data_HepG2like_SNR_L=0.0054699_diff=1.6188e-05.txt</td> <td>Synthetic gene expression knockdown (shRNA-seq) data immitating the ENCODE data for HepG2 cell line. Data size: 232 RBPs vs 464 experiments (2 replicates). SNR_L is the value of signal to noise ratio. Difference (diff) value tells the difference between replicate correlation coefficients of real and synthetic ENCODE data. Columns represent experiments, rows represent genes.</td> </tr> <tr> <td>data_K562like_SNR_L=0.0028692_diff=0.00017339.txt</td> <td>Synthetic gene expression knockdown (shRNA-seq) data immitating the ENCODE data for K562 cell line. Data size: 232 RBPs vs 464 experiments (2 replicates). SNR_L is the value of signal to noise ratio. Difference (diff) value tells the difference between replicate correlation coefficients of real and synthetic ENCODE data. Columns represent experiments, rows represent genes.</td> </tr> <tr> <td>network_HEPG2like_sparsity4.txt</td> <td>Synthetic scale-free gene regulatory network compatibile with data_HepG2like_SNR_L=0.0054699_diff=1.6188e-05.txt. Sparsity (average node degree) is 4 including selfloops. Direction should be read from columns to rows.</td> </tr> <tr> <td>network_K562like_sparsity4.txt</td> <td>Synthetic scale-free gene regulatory network compatibile with data_K562like_SNR_L=0.0028692_diff=0.00017339.txt. Sparsity (average node degree) is 4 including selfloops. Direction should be read from columns to rows.</td> </tr> <tr> <td>perturbations_HepG2&K562_2replicates.txt</td> <td>Perturbation matrix including information about knockeddown RBPs. Data size: 232 RBPs vs 464 experiments (2 replicates).</td> </tr> </tbody> </table> <p> </p> <p>Created by Garbulowski et al. (2024) as a part of the work entitled "Comprehensive analysis of the RBP regulome reveals functional modules and drug candidates in liver cancer"</p>
The VAROS Synthetic Underwater Data Set: Towards realistic multi-sensor underwater data with ground truth
<p>Underwater visual perception requires being able to deal with bad and rapidly varying illumination and with reduced visibility due to water turbidity. The verification of such algorithms is crucial for safe and efficient underwater exploration and intervention operations. Ground truth data play an important role in evaluating vision algorithms. However, obtaining ground truth from real underwater environments is in general very hard, if possible at all. In a synthetic underwater 3D environment, however, (nearly) all parameters are known and controllable, and ground truth data can be absolutely accurate in terms of geometry. In this paper, we present the VAROS environment, our approach to generating highly realistic underwater video and auxiliary sensor data with precise ground truth, built around the Blender modeling and rendering environment. VAROS allows for physically realistic motion of the simulated underwater (UW) vehicle including moving illumination. Pose sequences are created by first defining way-points for the simulated underwater vehicle which are expanded into a smooth vehicle course sampled at IMU data rate (200Hz). This expansion uses a vehicle dynamics model and a discrete-time controller algorithm that simulates the sequential following of the way-points. The scenes are rendered using the raytracing method, which generates realistic images, integrating direct light, and indirect volumetric scattering. The VAROS dataset version 1 provides images, inertial measurement unit (IMU) and depth gauge data, as well as ground truth poses, depth images and surface normal images.</p>
Synthetic Data of Transactions for Inmediate Loans' Fraud
<p><strong>This dataset contains realistic synthetic data generated with a commercial tool, taking as an input a real dataset of CaixaBank’s express loans for a timespan of 18 months. The real dataset was tagged in order to identify the confirmed and tentative fraud cases in which a fraudster has impersonate the client to claim that type of loan and steal client’s funds. The dataset includes several indicators that help fraud analysts to identify any suspicious behaviour of the user that could imply an impersonation or misbehaviour. This dataset was used in INFINITECH H2020 project to build an AI model for cyberfraud prevention in this type of operations, which are especially critical because of two factors. First, it is type of loan, an operation in which the fraudster can steal money that the client does not really own, so it can be stolen even from clients without funds on their accounts. Second, it is an operation that was offered to the clients to speed up the process of acquiring loans of small amounts. The fraudsters can take profit of that and proceed faster as well stealing that money. The detail of the data fields included in the dataset is specified in the table below.</strong></p> <p> </p> <table> <tbody> <tr> <td> <p><strong>Field name</strong></p> </td> <td> <p><strong>Value example</strong></p> </td> <td> <p><strong>Field description</strong></p> </td> </tr> <tr> <td> <p><strong>Fraud</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicates if a fraud was produced in the operation. (0 No; 1 Intent of fraud; 2 Completed fraud -money stolen-)</strong></p> </td> </tr> <tr> <td> <p><strong>PK_ANYOMES </strong></p> </td> <td> <p><strong>202102</strong></p> </td> <td> <p><strong>Year and month of the loan constitution operation</strong></p> </td> </tr> <tr> <td> <p><strong>PK_ANYOMESDIA </strong></p> </td> <td> <p><strong>20210207</strong></p> </td> <td> <p><strong>Day of the loan constitution operation</strong></p> </td> </tr> <tr> <td> <p><strong>PK_TSINSERCION </strong></p> </td> <td> <p><strong>06:28,0</strong></p> </td> <td> <p><strong>Time of the loan constitution operation</strong></p> </td> </tr> <tr> <td> <p><strong>IDE_USUCLO_ORIG </strong></p> </td> <td> <p><strong>1321946400</strong></p> </td> <td> <p><strong>User associated with the online banking contract and the client. It is an internal user ID which is used jointly with PK_CONTRATO to access the services under the online banking contract.</strong></p> </td> </tr> <tr> <td> <p><strong>PK_CONTRATO </strong></p> </td> <td> <p><strong>1096097250023219464</strong></p> </td> <td> <p><strong>online banking contract code. It is the identifier of the online banking services.</strong></p> </td> </tr> <tr> <td> <p><strong>FK_NUMPERSO </strong></p> </td> <td> <p><strong>27388223</strong></p> </td> <td> <p><strong>Unique ID that identifies the physical person (client) who is connecting to online banking</strong></p> </td> </tr> <tr> <td> <p><strong>IDE_SAU </strong></p> </td> <td> <p><strong>08875268</strong></p> </td> <td> <p><strong>Identifier used by the client to access online banking. This identifier is used jointly with CARPETA id to access online banking services. </strong></p> </td> </tr> <tr> <td> <p><strong>CARPETA </strong></p> </td> <td> <p><strong>49830679</strong></p> </td> <td> <p><strong>Folder the online banking services of the clients are stored. It is used jointly with the client's online banking identifier (ID_SAU).</strong></p> </td> </tr> <tr> <td> <p><strong>FK_COD_OPERACION </strong></p> </td> <td> <p><strong>03693</strong></p> </td> <td> <p><strong>Loan constitution transaction code. Unique ID that identifies the loan.</strong></p> </td> </tr> <tr> <td> <p><strong>DES_OPERACION </strong></p> </td> <td> <p><strong>CONSTITUCION PRESTAMO</strong></p> </td> <td> <p><strong>Description of the loan constitution operation.</strong></p> </td> </tr> <tr> <td> <p><strong>IP_TERMINAL </strong></p> </td> <td> <p><strong>AAHUAWPOTLXYxgaNLC zWp70Yp+MaW2i1qEkh0o=</strong></p> </td> <td> <p><strong>IP of the terminal or hash of the mobile device from which the client connects to online banking.</strong></p> </td> </tr> <tr> <td> <p><strong>FK_NUMPERSO_TIT_LOE </strong></p> </td> <td> <p><strong>27388223</strong></p> </td> <td> <p><strong>Identifier of the physical person that is the online banking contract holder. It can be different to FK_NUMPERSO, if FK_NUMPERSO is an authorised person to operate the online banking services of FK_NUMPERSO_TIT_LOE. It can happen both for FK_NUMPERSO_TIT_LOE representing physical or legal persons (enterprises).</strong></p> </td> </tr> <tr> <td> <p><strong>FK_CONTRATO_PPAL_OPE </strong></p> </td> <td> <p><strong>1001037520210005473</strong></p> </td> <td> <p><strong>Contract code of the savings account in which the loan is deposited. This is not the same contract as the online banking contract.</strong></p> </td> </tr> <tr> <td> <p><strong>FK_IMPORTE_PRINCIPAL </strong></p> </td> <td> <p><strong>1500</strong></p> </td> <td> <p><strong>Loan amount demanded.</strong></p> </td> </tr> <tr> <td> <p><strong>IND_MFA_OPE</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicator of the response of the SCA (Strong Customer Authentication) request decision algorithm for the loan consolidation operation. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>MESSAGE_MFA_OPE</strong></p> </td> <td> <p><strong>Konline bankingN USER AND DEVICE</strong></p> </td> <td> <p><strong>SCA (Strong Customer Authentication) request decision algorithm response message for loan consolidation operation.</strong></p> </td> </tr> <tr> <td> <p><strong>SALDO_ANTES_PRESTAMO</strong></p> </td> <td> <p><strong>100</strong></p> </td> <td> <p><strong>Balance of the account into which the loan is deposited just before the loan.</strong></p> </td> </tr> <tr> <td> <p><strong>POSICION_GLOBAL_ANTES_PRESTAMO</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Global balance of the client before the loan. (1: <1000; 2: 1000-10000; 3: 10000-50000; 4: 50000-250000; 5: >250000; -2: Data not found)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_NUEVO_IDE_SAU</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>If the identifier used to access online banking has been created in the last 48 hours. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>FECHA_ALTA_CLIENTE</strong></p> </td> <td> <p><strong>39246</strong></p> </td> <td> <p><strong>Indicate the date of registration with CaixaBank as a customer. When the physical person (FK_NUMPERSO) became a client of CaixaBank</strong></p> </td> </tr> <tr> <td> <p><strong>IND_ALTA_SIGN</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicates if the client has registered a sign in the last 48 hours. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_GMP_ANT</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicates if there has been a new primary mobile assignment in the 48 hours prior to the loan. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_INGRESO_NOMINA</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Indicate if the payroll of FK_NUMPERSO is domiciled at CaixaBank. (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_PENSION</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO has the pension domiciled in CaixaBank. (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_IMAGIN_BANK</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO is ImaginBank customer (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_EXTRANJERO</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO is a foreigner (0 National; 1 Foreigner)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_RESIDENTE</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO resides in Spain (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>FK_TIPREL</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Type of the ownership of the savings account in which the loan is deposited (values between 1 and 48). 1 means it is an account holder. Other values mean other type of relationships (i.e. "authorized person but not an owner of the account").</strong></p> </td> </tr> <tr> <td> <p><strong>FK_ORDREL</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Order of the ownership relationship. If there are more than account holder, in which position is the FK_NUMPERSO.</strong></p> </td> </tr> </tbody> </table> <p> </p>
I-BiDaaS - TID - Synthetic Mobility Data
<p>This is a synthetic data stream based on real-time, cell network events. These events are picked up by the antennas that are closer to the mobile phone thus providing an approximate location of the device. Every transaction of a mobile phone generates one of those events. A transaction can be, for instance, placing or receiving a call, sending or receiving an SMS, asking for a specific URL in your mobile phone browser, or sending a text message or a data transaction from/to any mobile phone app. There are also some synchronization events like, for instance, turning your mobile phone on or off, or when switching between location area networks (relatively big geographical areas comprising several cell towers).</p>
Synthetic data starch potato system Veenkoloniën
<p>Synthetic generated with a simulation model for the starch potato production systems in the Veenkoloniën</p>
Synthetic data for the Bourbonnais
<p>Synthetic data produced using a system dynamics model for beef production systems in the Bourbonnais France.</p>
Synthetic Data for Uplift Modeling and Heterogenous Treatment Effect with Known Counterfactuals and ITE
<p>This dataset is designed and simulated for evaluating uplift modeling. The data generation process is based on a logistic regression model - no real data is included or used for generating this dataset.</p> <p>This dataset has several signatures:</p> <ul> <li>It generates features with various patterns associated with the outcome variable and the causal effect (or treatment effect). Thus it is suitable for evaluating feature importance and model interpretation for uplift modeling.</li> <li>The true counterfactual outcomes under control and treatment are known for each user, as well as the true ITE (Individual treatment effect).</li> </ul> <p>This dataset consists of 50 trials (replicates with different random seeds), each trial with 20,000 samples and 36 features. The outcome variable is binary, which makes this dataset for classification problems. The samples are equally split for the control and treatment groups (10,000 samples in each group in each trial).</p> <p>The generated data has three types of features: (1) uplift features influencing the treatment effect on the conversion probability; (2) classification features affecting the conversion probability but independent of the treatment effect; and (3) irrelevant features that are independent of both conversion probability and the treatment effect.</p> <p>To simulate the relationship between uplift features and the treatment effect and classification features and outcome probability, we implement six types of association patterns in the data generation process: linear, quadratic, cubic, ReLU (Rectified Linear Unit), trigonometric function sine, and cosine.</p> <p>In this data set, there are 36 features in total, including 10 classification features, 6 uplift features, and 20 irrelevant features.</p> <p>Column names:</p> <p> Trial ID: 'trial_id'<br> Experiment group label: 'treatment_group_key'<br> Outcome variable (classification label): 'conversion'<br> Feature names: ['x1_informative',<br> 'x2_informative',<br> 'x3_informative',<br> 'x4_informative',<br> 'x5_informative',<br> 'x6_informative',<br> 'x7_informative',<br> 'x8_informative',<br> 'x9_informative',<br> 'x10_informative',<br> 'x11_irrelevant',<br> 'x12_irrelevant',<br> 'x13_irrelevant',<br> 'x14_irrelevant',<br> 'x15_irrelevant',<br> 'x16_irrelevant',<br> 'x17_irrelevant',<br> 'x18_irrelevant',<br> 'x19_irrelevant',<br> 'x20_irrelevant',<br> 'x21_irrelevant',<br> 'x22_irrelevant',<br> 'x23_irrelevant',<br> 'x24_irrelevant',<br> 'x25_irrelevant',<br> 'x26_irrelevant',<br> 'x27_irrelevant',<br> 'x28_irrelevant',<br> 'x29_irrelevant',<br> 'x30_irrelevant',<br> 'x31_uplift_increase',<br> 'x32_uplift_increase',<br> 'x33_uplift_increase',<br> 'x34_uplift_increase',<br> 'x35_uplift_increase',<br> 'x36_uplift_increase']<br> True underlying control conversion probability: 'control_conversion_prob'<br> True underlying treatment conversion probability: 'treatment1_conversion_prob'<br> True treatment effect: 'treatment1_true_effect'</p>
Data for manuscript "Creating boundaries along a synthetic frequency dimension" in Nature Communications
<p>Data for manuscript "Creating boundaries along a synthetic frequency dimension"</p> <p>https://www.nature.com/articles/s41467-022-31140-7</p> <p>https://arxiv.org/abs/2203.11296</p>
Synthetic 3D PPGIS Data _ Turku - Finland
<p>3D PPGIS data generated synthetically in Turku, Finland. Data is created in an approximately 2 km<sup>2 </sup>area near center. 150X150 m grid cells were used for data generation.</p>
Synthetic geospatial data for performance analysis of geospatial database systems
<p>This dataset contains a set of synthetic data that can be used to evaluate the efficiency of geosaptial datasbases. </p> <p>The datasets is composed of four json file, characterized by different size. They can be used to analyze the scalability of geospatial datasets with respect to the database size.</p> <p>Each json file contains a set of "points", each one characterized by a set of random attributes (description, url of a picture linked to the point, creation date, delete date, update date, identifier, partition identifier).</p> <p>The synthetically generated points are uniformly distributed among the world.</p>
Source data for "Synthetic gauge fields for phonon transport in a nano-optomechanical system"
<ul> <li>Experimental raw data for density plots in Fig 2. Each .csv contains an array, where 1st row corresponds to x_axis (mechanical frequency in MHz for panels 1,2,3,4) and first column the y_axis (optical frequency in THz for panel 1, modulation frequency in MHz for panels 2,3,4). First nonzero component is the 2nd for each array. Remaining array elements contain the z values (Thermomechanical noise spectral for panel 1, Amplitude of driven responses for panels 2,3,4). An illustrative example of plotting in an ipython notebook follows:</li> </ul> <p> %pylab inline</p> <p> A= genfromtxt('Fig2_data_modVolt=0mV_experiment.csv', delimiter=',') </p> <p> x = A[0,1:]<br> y = A[1:,0]<br> z = A[1:,1:]<br> imshow(z,aspect='auto',vmin=z.min(),vmax=z.max(),extent=[x.min(),x.max(),y.min(),y.max()],cmap='magma') </p> <ul> <li> Theoretical data for panel 4 in Fig 2, stored in a .csv with the same structure as previous.</li> <li> Raw experimental data for upper panels in Fig 3. Each .csv contains an array where 1st row corresponds to x_axis (modulation phase) and first column the y_axis (optical frequency in THz). Z values contain the experimental signal proportional to the Y optical quadrature of the transferred mode.</li> <li>Theoretical data for lower panels in Fig 3, stored in a .csv with the same structure as previous.</li> <li>Jupyter notebook to produce and plot typical data for Fig 4: phononic amplitude averaged over 100 disorder realizations, normalized to the maximum value (*extra_dependencies: Kwant Python library: <a href="https://kwant-project.org/">https://kwant-project.org/</a>).</li> </ul>
Source data for Synthetic dynamic hydrogels promote degradation-independent in vitro organogenesis
<p>Source data and statistical analysis results for Synthetic dynamic hydrogels promote degradation-independent in vitro organogenesis</p>
Processed Synthetic Real-World Data for tristate modelling
<p>This model learning dataset is created out of the <a href="https://zenodo.org/record/7409763">Raw Synthetic RWD</a> raw dataset, including some of the original attributes. It is distributed in JOBLIB files, where .joblib files contain the vectors and _ids.joblib contain the ID of the person from which each vector is extracted.</p> <p>This is useful in case it is needed to map the vectors to metadata about the people that are found in the original raw dataset. Note that corresponds to , or , depending on the dataset.</p> <p>The split is roughly 60% of the people are in the training dataset, and 20% in each of the validation and the testing datasets. The input attributes are the age, the short-term averages and the trends of the current week’s BMI, steps walked, calories burned, sleep quality, mood and water consumption, as well as the previous week’s short-term average and trend of the answer to the health self-assessment question.</p> <p>The outcome to be predicted is a tristate quantized version of the health self-assessment answer to be given in the current week. The dataset is normalized based on the training set. The means and standard deviations used can be found in the train_statistics.joblib file. Finally, the output_descriptions.joblib file contains descriptions of the outcomes to be predicted (not actually needed, since included here).</p>
Processed Synthetic Real-World Data for binary modelling
<p>This model learning dataset is created out of the <a href="https://zenodo.org/record/7409763">Raw Synthetic RWD</a> raw dataset, including some of the original attributes. It is distributed in JOBLIB files, where .joblib files contain the vectors and _ids.joblib contain the ID of the person from which each vector is extracted.</p> <p>This is useful in case it is needed to map the vectors to metadata about the people that are found in the original raw dataset. Note that corresponds to , or , depending on the dataset. The split is roughly 60% of the people are in the training dataset, and 20% in each of the validation and the testing datasets. The input attributes are the age, the short-term averages and the trends of the current week’s BMI, steps walked, calories burned, sleep quality, mood and water consumption, as well as the previous week’s short-term average and trend of the answer to the health self-assessment question.</p> <p>The outcome to be predicted is the binary quantized health self-assessment answer to be given in the current week. The dataset is normalized based on the training set. The means and standard deviations used can be found in the train_statistics.joblib file. Finally, the output_descriptions.joblib file contains descriptions of the outcomes to be predicted (not actually needed, since included here).</p>
Raw Synthetic Real-World Data
<p>Real World Data from a simulator for 1,000 people belonging to any of four behavioural groups: athletic, normal, unfit and feeble, simulated over 2 years and 3 months.</p> <p>The dataset is organised into 4 CSV files:</p> <ul> <li>aggregations.csv contains the daily physiological data, including their short- & long-term averages and their trends. Every row corresponds to a day for a person.</li> <li>answers.csv contains the answers to questions. Each row corresponds to a question answered.</li> <li>activity_summaries.csv contains summaries of the detected activities. Each row corresponds to an activity someone performed.</li> <li>participants_features.csv contains the demographic info of each simulated person. This is static info, every row corresponding to one person. An ML engineer can combine information from any of these files to setup model learning datasets. This is done in the “Processed Synthetic RWD for binary modelling” and the “Processed Synthetic RWD for tristate modelling” datasets that are also shared here.</li> </ul>
Open synthetic data on travel and charging demand of battery electric cars: An agent-based simulation on three charging behavior archetypes
<p><strong>Background</strong></p> <p>Battery electric vehicles (BEVs) are crucial for a sustainable transportation system. As more people adopt BEVs, it becomes increasingly important to accurately assess the demand for charging infrastructure. However, much of the current research on charging infrastructure relies on outdated assumptions, such as the assumption that all BEV owners have access to home chargers and the "Liquid-fuel" mental model. To address this issue, we simulate the travel and charging demand on three charging behavior archetypes. We use a large synthetic population of Sweden, including detailed individual characteristics, such as dwelling types (detached house vs. apartment) and activity plans (for an average weekday). This data repository aims to provide the BEV simulation's input, assumptions, and output so that other studies can use them to study sizing and location design of charging infrastructure, grid impact, etc.</p> <p>A journal paper published in Transportation Research Part D: Transport and Environment details the method to create the data (particularly Section 2.2 BEV simulation).</p> <p><a href="https://doi.org/10.1016/j.trd.2023.103645">https://doi.org/10.1016/j.trd.2023.103645</a></p> <p><strong>Methodology</strong></p> <p>This data product is centered on the 1.7 million inhabitants of the Västra Götaland (VG) region, which includes the second largest city in Sweden, Gothenburg. We specifically simulated 284,000 car agents who live in VG, representing 35% of all car users and 18% of the total population in the region. They spend their simulation day (representing an average weekday) in a variety of locations throughout Sweden.</p> <p>This open data repository contains the core model inputs and outputs. The numbers in parentheses correspond to the data sets. We use individual agents' activity plans (1) and travel trajectories from MATSim simulation for the BEV simulation (2), in which we consider overnight charger access (3), car fleet composition referencing the current private car fleet in Sweden (4), and Swedish road network with slope information (5) with realistic BEV charging & discharging dynamics. For the BEV simulation, we tested ten scenarios of charging behavior archetypes and fast charging powers (6). The output includes the time history of travel trajectories and charging of the simulated BEVs across the different scenarios (7).</p> <p><strong>Data description</strong></p> <p>The current data product covers seven data files.</p> <p><strong>(1) Agents' experienced activity plans</strong></p> <p>File name: 1_activity_plans.csv</p> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>person</p> </td> <td> <p>Agent ID</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>act_id</p> </td> <td> <p>Activity index of each agent</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>deso</p> </td> <td> <p>Zone code of Demographic statistical areas (DeSO)<sup>1</sup></p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>POINT_X</p> </td> <td> <p>Coordinate X of activity location (SWEREF99TM)</p> </td> <td> <p>Float</p> </td> <td> <p>meter</p> </td> </tr> <tr> <td> <p>POINT_Y</p> </td> <td> <p>Coordinate Y of activity location (SWEREF99TM)</p> </td> <td> <p>Float</p> </td> <td> <p>meter</p> </td> </tr> <tr> <td> <p>act_purpose</p> </td> <td> <p>Activity purpose (work, home, other)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>mode</p> </td> <td> <p>Transport mode to reach the activity location (car)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>dep_time</p> </td> <td> <p>Departure time in decimal hour (0-23.99)</p> </td> <td> <p>Float</p> </td> <td> <p>hour</p> </td> </tr> <tr> <td> <p>trav_time</p> </td> <td> <p>Travel time to reach the activity location</p> </td> <td> <p>String</p> </td> <td> <p>hour:minute:second</p> </td> </tr> <tr> <td> <p>trav_time_min</p> </td> <td> <p>Travel time in decimal minute</p> </td> <td> <p>Float</p> </td> <td> <p>minute</p> </td> </tr> <tr> <td> <p>speed</p> </td> <td> <p>Travel speed to reach the activity location</p> </td> <td> <p>Float</p> </td> <td> <p>km/h</p> </td> </tr> <tr> <td> <p>distance</p> </td> <td> <p>Travel distance between the origin and the destination</p> </td> <td> <p>Float</p> </td> <td> <p>km</p> </td> </tr> <tr> <td> <p>act_start</p> </td> <td> <p>Start time of activity in minute (0-1439)</p> </td> <td> <p>Integer</p> </td> <td> <p>minute</p> </td> </tr> <tr> <td> <p>act_time</p> </td> <td> <p>Activity duration in decimal minute</p> </td> <td> <p>Float</p> </td> <td> <p>minute</p> </td> </tr> <tr> <td> <p>act_end</p> </td> <td> <p>End time of activity in decimal hour (0-23.99)</p> </td> <td> <p>Float</p> </td> <td> <p>hour</p> </td> </tr> <tr> <td> <p>score</p> </td> <td> <p>Utility score of the simulation day given by MATSim</p> </td> <td> <p>Float</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> <p>1 <a href="https://www.scb.se/vara-tjanster/oppna-data/oppna-geodata/deso--demografiska-statistikomraden/">https://www.scb.se/vara-tjanster/oppna-data/oppna-geodata/deso--demografiska-statistikomraden/</a></p> <p> </p> <p><strong>(2) Travel trajectories</strong></p> <p>File name: 2_input_zip</p> <p>Produced by MATSim simulation, the zip folder contains ten files (events_batch_X.csv.gz, X=1, 2, …, 10) of input events for the BEV simulation. They are the moving trajectories of the car agents in their simulation days.</p> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>time</p> </td> <td> <p>Time in second in a simulation day (0-86399)</p> </td> <td> <p>Integer</p> </td> <td> <p>Second</p> </td> </tr> <tr> <td> <p>type</p> </td> <td> <p>Event type defined by MATSim simulation<sup>2</sup></p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>person</p> </td> <td> <p>Agent ID</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>link</p> </td> <td> <p>Nearest road link consistent with (5)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>vehicle</p> </td> <td> <p>Vehicle ID identical to person</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> <p><sup>2 </sup>One typical episode of MATSim simulation events: Activity ends (actend) -> Agent’s vehicle enters traffic (vehicle enters traffic) -> Agent’s vehicle moves from previous road segment to its next connected one (left link) -> Agent’s vehicle leaves traffic for activity (vehicle leaves traffic) -> Activity starts (actstart)</p> <p> </p> <p><strong>(3) Overnight charger access</strong></p> <p>File name: 3_home_charger_access.csv</p> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>person</p> </td> <td> <p>Agent ID</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>home_charger</p> </td> <td> <p>Whether an agent has access to a home garage charger/living in a detached house (0=no, 1=yes)</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> <p> </p> <p><strong>(4) Car fleet composition</strong></p> <p>File name: 4_car_fleet.csv</p> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>person</p> </td> <td> <p>Agent ID</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>income_class</p> </td> <td> <p>Income group (0=None, 1=below 180K, 2=180K-300K, 3=300K-420K, 4=above 420K)</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>car</p> </td> <td> <p>Car model class (B=40 kWh, C=60 kWh, D=100 kWh)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> <p> </p> <p>(<strong>5) Road network with slope information</strong></p> <p>File name: 5_road_network_with_slope.shp (5 files in total)</p> <table> <tbody> <tr> <td> <p>Column</p> </td> <td> <p>Description</p> </td> <td> <p>Data type</p> </td> <td> <p>Unit</p> </td> </tr> <tr> <td> <p>length</p> </td> <td> <p>The length of road link</p> </td> <td> <p>Float</p> </td> <td> <p>meter</p> </td> </tr> <tr> <td> <p>freespeed</p> </td> <td> <p>Free speed</p> </td> <td> <p>Float</p> </td> <td> <p>km/h</p> </td> </tr> <tr> <td> <p>capacity</p> </td> <td> <p>Number of vehicles</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>permlanes</p> </td> <td> <p>Number of lanes</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>oneway</p> </td> <td> <p>Whether the segment is one-way (0=no, 1=yes)</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>modes</p> </td> <td> <p>Transport mode (car)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>link_id</p> </td> <td> <p>Link ID</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>from_node</p> </td> <td> <p>Start node of the link</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>to_node</p> </td> <td> <p>End node of the link</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>count</p> </td> <td> <p>Aggregated traffic (number of cars travelled per day)</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>slope</p> </td> <td> <p>Slope in percent from -6% to 6%</p> </td> <td> <p>Float</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>geometry</p> </td> <td> <p>LINESTRING (SWEREF99TM)</p> </td> <td> <p>geometry</p> </td> <td> <p>meter</p> </td> </tr> </tbody> </table> <p> </p> <p><strong>(6) Simulation scenarios specifying the parameter sets</strong></p> <p>File name: 6_scenarios.txt</p> <table> <tbody> <tr> <td> <p><strong>Parameter set</strong></p> <p><strong>(paraset)</strong></p> </td> <td> <p><strong>Strategy 1</strong></p> </td> <td> <p><strong>Strategy 2</strong></p> </td> <td> <p><strong>Strategy 3</strong></p> </td> <td> <p><strong>Fast charging power (kW)</strong></p> </td> <td> <p><strong>Minimum parking time for charging (min)</strong></p> </td> <td> <p><strong>Intermediate charging power (kW)</strong></p> </td> </tr> <tr> <td> <p>0</p> </td> <td> <p>0.2</p> </td> <td> <p>0.2</p> </td> <td> <p>0.9</p> </td> <td> <p>150</p> </td> <td> <p>5</p> </td> <td> <p>22</p> </td> </tr> <tr> <td> <p>1</p> </td> <td> <p>0.2</p> </td> <td> <p>0.2</p> </td> <td> <p>0.9</p> </td> <td> <p>50</p> </td> <td> <p>5</p> </td> <td> <p>22</p> </td> </tr> <tr> <td> <p>2</p> </td> <td> <p>0.3</p> </td> <td> <p>0.3</p> </td> <td> <p>0.9</p> </td> <td> <p>150</p> </td> <td> <p>5</p> </td> <td> <p>22</p> </td> </tr> <tr> <td> <p>3</p> </td> <td> <p>0.3</p> </td> <td> <p>0.3</p> </td> <td> <p>0.9</p> </td> <td> <p>50</p> </td> <td> <p>5</p> </td> <td> <p>22</p> </td> </tr> </tbody> </table> <p> </p> <p><strong>(7) Time history of travel trajectories and charging of the simulated BEVs</strong></p> <p>File name: 7_output.zip</p> <p>Produced by the BEV simulation, the zip folder contains four files (parasetX.csv.gz, X=1, 2, 3, 4) corresponding to the four parameter sets specified in (6). They are the moving trajectories of the car agents with simulated energy and charging time history in their simulation days.</p> <table> <tbody> <tr> <td> <p><strong>Column</strong></p> </td> <td> <p><strong>Description</strong></p> </td> <td> <p><strong>Data type</strong></p> </td> <td> <p><strong>Unit</strong></p> </td> </tr> <tr> <td> <p>person</p> </td> <td> <p>Agent ID</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>home_charger</p> </td> <td> <p>Whether an agent has access to a home garage charger/living in a detached house (0=no, 1=yes)</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>car</p> </td> <td> <p>Car model class (B=40 kWh, C=60 kWh, D=100 kWh)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>seq</p> </td> <td> <p>Sequence ID of time history by agent</p> </td> <td> <p>Integer</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>time</p> </td> <td> <p>Time (0-86399)</p> </td> <td> <p>Integer</p> </td> <td> <p>Second</p> </td> </tr> <tr> <td> <p>purpose</p> </td> <td> <p>Valid for activities (home, work, school, other)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>type</p> </td> <td> <p>Event type defined by MATSim simulation</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>link</p> </td> <td> <p>Link ID (link_id in File 5)</p> </td> <td> <p>String</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>distance_driven</p> </td> <td> <p>Cumulative driven distance in the simulation day</p> </td> <td> <p>Float</p> </td> <td> <p>km</p> </td> </tr> <tr> <td> <p>energy_1</p> </td> <td> <p>Energy consumed while driving (-) or charging (+) (Strategy 1)</p> </td> <td> <p>Float</p> </td> <td> <p>kWh</p> </td> </tr> <tr> <td> <p>energy_2</p> </td> <td> <p>Energy consumed while driving (-) or charging (+) (Strategy 2)</p> </td> <td> <p>Float</p> </td> <td> <p>kWh</p> </td> </tr> <tr> <td> <p>energy_3</p> </td> <td> <p>Energy consumed while driving (-) or charging (+) (Strategy 3)</p> </td> <td> <p>Float</p> </td> <td> <p>kWh</p> </td> </tr> <tr> <td> <p>charger_1</p> </td> <td> <p>Power rating of the charger (Strategy 1)</p> </td> <td> <p>Float</p> </td> <td> <p>kW</p> </td> </tr> <tr> <td> <p>charger_2</p> </td> <td> <p>Power rating of the charger (Strategy 2)</p> </td> <td> <p>Float</p> </td> <td> <p>kW</p> </td> </tr> <tr> <td> <p>charger_3</p> </td> <td> <p>Power rating of the charger (Strategy 3)</p> </td> <td> <p>Float</p> </td> <td> <p>kW</p> </td> </tr> <tr> <td> <p>soc_1</p> </td> <td> <p>State of charge (0-1, Strategy 1)</p> </td> <td> <p>Float</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>soc_2</p> </td> <td> <p>State of charge (0-1, Strategy 2)</p> </td> <td> <p>Float</p> </td> <td> <p>-</p> </td> </tr> <tr> <td> <p>soc_3</p> </td> <td> <p>State of charge (0-1, Strategy 3)</p> </td> <td> <p>Float</p> </td> <td> <p>-</p> </td> </tr> </tbody> </table> <p> </p>
Synthetic parcel data for Lyon
<p>This data set contains a synthetic parcel data set for Lyon that has been generated during the Horizon 2020 project LEAD. It represents individual parcels that need to be delivered during an average day in Lyon with individual coordinates, based on a synthetic population for the city.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.