Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,754
datasets available to search
ShareScore release 0.7.1
Dataset results
1,754 results for “synthetic”
Synthetic dataset accompanying Neural Image Compression for Gigapixel Histopathology Image Analysis
<p>This dataset was used to develop and evaluate the main method proposed in the paper "Neural Image Compression for Gigapixel Histopathology Image Analysis" published in IEEE Transactions on Pattern Analysis and Machine Intelligence with DOI 10.1109/TPAMI.2019.2936841. Please refer to the paper for a detailed description of the dataset.</p> <p>The dataset consists of a set of 50000 images and 50000 associated ground truth masks, distributed into training and test partitions. The name of each file follows the pattern "{id}_{tilted_label}_{nontilted_label}_{tilted_size}_{nontilted_size}_{kind}.png" where:<br> * id: unique identifier within each partition.<br> * tilted_label: image-level label corresponding to the tilted rectangle.<br> * nontilted_label: image-level label corresponding to the non-tilted rectangle.<br> * tilted_size: longest size of the tilted rectangle.<br> * nontilted_size: longest size of the non-tilted rectangle.<br> * kind: either "tile" or "mask" image type.</p> <p>The images are distributed into several data partitions used during cross-validation and fully described in "mnist_folds_set.json". Please rename "mnist_folds_set.json.removethis" into "mnist_folds_set.json".</p> <p>The code to recreate this dataset can be found in https://github.com/davidtellez/neural-image-compression.</p>
Synthetic Collision Dataset for Spacecraft Collision Avoidance
<p>This dataset is intended to be used as a banchmark for testing collision avoidance strategies.</p> <p>It is made of 21000000 relative geometries between LEO space objects, 1000 of which are true collision (miss-distance smaller than combined hard body radius).<br>These relative geometries are expressed as target and chaser 6-dimensional state vectors (cartesian coordinates) at time of closest approach.</p> <p>The relative geometry of the encounters are statistically matched to the ESA's Kelvins dataset for the collision avoidance challenge through statistical fitting methods.</p> <p>The collision proportion is tuned to reflect a 1year mission in LEO orbit with an a-priori collision probability of 1e-3 (yearly) and a 21 collision warnings per year.</p> <p>*<em><strong> Implementation Description *</strong></em></p> <p>The dataset is made of a series of .mat files storing the following variables:</p> <div> <ul> <li>'rv_t', target's cartesian state at TCA (km, km/s, in ECI) 6xN vector</li> <li>'rv_c', chaser's cartesian state at TCA (km, km/s, in ECI) 6xN vector </li> <li>'Ct', target's position covariance matrix at TCA (km^2, in target's RTN at TCA) 3x3xN </li> <li>'Cc', chaser's position covariance matrix at TCA (km^2, in chaser's RTN at TCA) 3x3xN </li> <li>'Rc', combined hard body radius (m) 1xN</li> <li>'CollFlag', logic value of collision 1xN (0: no-collision, 1: collision)</li> <li>'missDistance', miss distance at TCA (km) Nx1</li> </ul> <p>The name of the .mat file is formatted as:</p> <p>batch_<batch start index>.mat</p> <p>Each batch file has a maximum dimension of N = 1e5.</p> </div>
Catalog of synthetic seismic records from mineral physics and travel-time tables from Waszek et al., 2021, Nature Geoscience
<p>This release is associated with the accepted publication in Nature Geoscience:</p> <p>Waszek L., Tauzin B., Schmerr N., Ballmer M. and Afonso J.C. A poorly mixed mantle transition zone and its thermal state inferred from seismic waves. Nature Geoscience, 2021.</p> <p>This dataset must be used in conjunction with the NoLimit software package (https://zenodo.org/record/5512805).</p> <p>Both the software and dataset allow the prediction of synthetic seismic waveforms for SS and PP-precursors from mineral physics models, as well as their processing for reconstructing the surface of seismic boundaries associated with major mineralogical phase transitions in the Earth’s mantle (namely, the 410-km and 660-km depth discontinuities).</p> <p>For technical reasons (storage and quick access), the catalog is downsampled with respect to the one in Waszek et al. (2021), and it is provided with the HDF5 format. For more advanced applications such as changing mantle composition, or generating waveforms for deeper earthquakes, please contact Benoit Tauzin (benoit.tauzin@univ-lyon1.fr) and Lauren Waszek (lauren.waszek@jcu.edu.au).</p> <p>The dataset includes:</p> <p>* A fixed mantle composition, which is a mechanical mixture of basalt and harzburgite with a fraction of basalt f=0.2.<br> * A downsampled catalog of synthetic waveforms for event depths between 0 and 80 km by step of 10 km (enough for reproducing the processing of observed SS and PP precursors waveforms).<br> * Adiabatic temperature gradients with potential temperature Tpot between 1200 and 2100 K by step of 100 K.</p> <p>This catalog and associated travel-time tables will allow any user to generate synthetic waveforms for any moment tensor, and events within the pre-defined depth interval.<br> </p> <p><strong>How to cite this material?</strong></p> <p>Any use of the datasets or software must refer to:</p> <p>The reference paper: Waszek L., Tauzin B., Schmerr N., Ballmer M., Afonso J.C. A poorly mixed mantle transition zone and its thermal state inferred from seismic waves. Nature Geoscience. 2021.<br> <br> Software: Tauzin, Benoit, & Waszek, Lauren. (2021). NoLiMit MATLAB package v1.0. Non-Linear Bayesian partition Modeling of the Earth's Mantle Transition zone (Version 1). Zenodo. https://doi.org/10.5281/zenodo.5512805<br> <br> Datasets: Tauzin, Benoit, Waszek, Lauren, & Afonso, Juan Carlos. (2021). Catalog of synthetic seismic records from mineral physics and travel-time tables from Waszek et al., 2021, Nature Geoscience (Version 1) [Data set]. Zenodo. https://doi.org/10.5281/zenodo.5512035</p>
The VAROS Synthetic Underwater Data Set: Towards realistic multi-sensor underwater data with ground truth
<p>Underwater visual perception requires being able to deal with bad and rapidly varying illumination and with reduced visibility due to water turbidity. The verification of such algorithms is crucial for safe and efficient underwater exploration and intervention operations. Ground truth data play an important role in evaluating vision algorithms. However, obtaining ground truth from real underwater environments is in general very hard, if possible at all. In a synthetic underwater 3D environment, however, (nearly) all parameters are known and controllable, and ground truth data can be absolutely accurate in terms of geometry. In this paper, we present the VAROS environment, our approach to generating highly realistic underwater video and auxiliary sensor data with precise ground truth, built around the Blender modeling and rendering environment. VAROS allows for physically realistic motion of the simulated underwater (UW) vehicle including moving illumination. Pose sequences are created by first defining way-points for the simulated underwater vehicle which are expanded into a smooth vehicle course sampled at IMU data rate (200Hz). This expansion uses a vehicle dynamics model and a discrete-time controller algorithm that simulates the sequential following of the way-points. The scenes are rendered using the raytracing method, which generates realistic images, integrating direct light, and indirect volumetric scattering. The VAROS dataset version 1 provides images, inertial measurement unit (IMU) and depth gauge data, as well as ground truth poses, depth images and surface normal images.</p>
Synthetic memory circuits for stable cell reprogramming in plants
<p>The data supporting the publication: Synthetic memory circuits for stable cell reprogramming in plants</p>
Synthetic Data of Transactions for Inmediate Loans' Fraud
<p><strong>This dataset contains realistic synthetic data generated with a commercial tool, taking as an input a real dataset of CaixaBank’s express loans for a timespan of 18 months. The real dataset was tagged in order to identify the confirmed and tentative fraud cases in which a fraudster has impersonate the client to claim that type of loan and steal client’s funds. The dataset includes several indicators that help fraud analysts to identify any suspicious behaviour of the user that could imply an impersonation or misbehaviour. This dataset was used in INFINITECH H2020 project to build an AI model for cyberfraud prevention in this type of operations, which are especially critical because of two factors. First, it is type of loan, an operation in which the fraudster can steal money that the client does not really own, so it can be stolen even from clients without funds on their accounts. Second, it is an operation that was offered to the clients to speed up the process of acquiring loans of small amounts. The fraudsters can take profit of that and proceed faster as well stealing that money. The detail of the data fields included in the dataset is specified in the table below.</strong></p> <p> </p> <table> <tbody> <tr> <td> <p><strong>Field name</strong></p> </td> <td> <p><strong>Value example</strong></p> </td> <td> <p><strong>Field description</strong></p> </td> </tr> <tr> <td> <p><strong>Fraud</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicates if a fraud was produced in the operation. (0 No; 1 Intent of fraud; 2 Completed fraud -money stolen-)</strong></p> </td> </tr> <tr> <td> <p><strong>PK_ANYOMES </strong></p> </td> <td> <p><strong>202102</strong></p> </td> <td> <p><strong>Year and month of the loan constitution operation</strong></p> </td> </tr> <tr> <td> <p><strong>PK_ANYOMESDIA </strong></p> </td> <td> <p><strong>20210207</strong></p> </td> <td> <p><strong>Day of the loan constitution operation</strong></p> </td> </tr> <tr> <td> <p><strong>PK_TSINSERCION </strong></p> </td> <td> <p><strong>06:28,0</strong></p> </td> <td> <p><strong>Time of the loan constitution operation</strong></p> </td> </tr> <tr> <td> <p><strong>IDE_USUCLO_ORIG </strong></p> </td> <td> <p><strong>1321946400</strong></p> </td> <td> <p><strong>User associated with the online banking contract and the client. It is an internal user ID which is used jointly with PK_CONTRATO to access the services under the online banking contract.</strong></p> </td> </tr> <tr> <td> <p><strong>PK_CONTRATO </strong></p> </td> <td> <p><strong>1096097250023219464</strong></p> </td> <td> <p><strong>online banking contract code. It is the identifier of the online banking services.</strong></p> </td> </tr> <tr> <td> <p><strong>FK_NUMPERSO </strong></p> </td> <td> <p><strong>27388223</strong></p> </td> <td> <p><strong>Unique ID that identifies the physical person (client) who is connecting to online banking</strong></p> </td> </tr> <tr> <td> <p><strong>IDE_SAU </strong></p> </td> <td> <p><strong>08875268</strong></p> </td> <td> <p><strong>Identifier used by the client to access online banking. This identifier is used jointly with CARPETA id to access online banking services. </strong></p> </td> </tr> <tr> <td> <p><strong>CARPETA </strong></p> </td> <td> <p><strong>49830679</strong></p> </td> <td> <p><strong>Folder the online banking services of the clients are stored. It is used jointly with the client's online banking identifier (ID_SAU).</strong></p> </td> </tr> <tr> <td> <p><strong>FK_COD_OPERACION </strong></p> </td> <td> <p><strong>03693</strong></p> </td> <td> <p><strong>Loan constitution transaction code. Unique ID that identifies the loan.</strong></p> </td> </tr> <tr> <td> <p><strong>DES_OPERACION </strong></p> </td> <td> <p><strong>CONSTITUCION PRESTAMO</strong></p> </td> <td> <p><strong>Description of the loan constitution operation.</strong></p> </td> </tr> <tr> <td> <p><strong>IP_TERMINAL </strong></p> </td> <td> <p><strong>AAHUAWPOTLXYxgaNLC zWp70Yp+MaW2i1qEkh0o=</strong></p> </td> <td> <p><strong>IP of the terminal or hash of the mobile device from which the client connects to online banking.</strong></p> </td> </tr> <tr> <td> <p><strong>FK_NUMPERSO_TIT_LOE </strong></p> </td> <td> <p><strong>27388223</strong></p> </td> <td> <p><strong>Identifier of the physical person that is the online banking contract holder. It can be different to FK_NUMPERSO, if FK_NUMPERSO is an authorised person to operate the online banking services of FK_NUMPERSO_TIT_LOE. It can happen both for FK_NUMPERSO_TIT_LOE representing physical or legal persons (enterprises).</strong></p> </td> </tr> <tr> <td> <p><strong>FK_CONTRATO_PPAL_OPE </strong></p> </td> <td> <p><strong>1001037520210005473</strong></p> </td> <td> <p><strong>Contract code of the savings account in which the loan is deposited. This is not the same contract as the online banking contract.</strong></p> </td> </tr> <tr> <td> <p><strong>FK_IMPORTE_PRINCIPAL </strong></p> </td> <td> <p><strong>1500</strong></p> </td> <td> <p><strong>Loan amount demanded.</strong></p> </td> </tr> <tr> <td> <p><strong>IND_MFA_OPE</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicator of the response of the SCA (Strong Customer Authentication) request decision algorithm for the loan consolidation operation. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>MESSAGE_MFA_OPE</strong></p> </td> <td> <p><strong>Konline bankingN USER AND DEVICE</strong></p> </td> <td> <p><strong>SCA (Strong Customer Authentication) request decision algorithm response message for loan consolidation operation.</strong></p> </td> </tr> <tr> <td> <p><strong>SALDO_ANTES_PRESTAMO</strong></p> </td> <td> <p><strong>100</strong></p> </td> <td> <p><strong>Balance of the account into which the loan is deposited just before the loan.</strong></p> </td> </tr> <tr> <td> <p><strong>POSICION_GLOBAL_ANTES_PRESTAMO</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Global balance of the client before the loan. (1: <1000; 2: 1000-10000; 3: 10000-50000; 4: 50000-250000; 5: >250000; -2: Data not found)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_NUEVO_IDE_SAU</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>If the identifier used to access online banking has been created in the last 48 hours. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>FECHA_ALTA_CLIENTE</strong></p> </td> <td> <p><strong>39246</strong></p> </td> <td> <p><strong>Indicate the date of registration with CaixaBank as a customer. When the physical person (FK_NUMPERSO) became a client of CaixaBank</strong></p> </td> </tr> <tr> <td> <p><strong>IND_ALTA_SIGN</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicates if the client has registered a sign in the last 48 hours. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_GMP_ANT</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicates if there has been a new primary mobile assignment in the 48 hours prior to the loan. (0 No; 1 Yes; -1 Unknown)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_INGRESO_NOMINA</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Indicate if the payroll of FK_NUMPERSO is domiciled at CaixaBank. (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_PENSION</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO has the pension domiciled in CaixaBank. (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_IMAGIN_BANK</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO is ImaginBank customer (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_EXTRANJERO</strong></p> </td> <td> <p><strong>0</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO is a foreigner (0 National; 1 Foreigner)</strong></p> </td> </tr> <tr> <td> <p><strong>IND_RESIDENTE</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Indicate if FK_NUMPERSO resides in Spain (0 No; 1 Yes)</strong></p> </td> </tr> <tr> <td> <p><strong>FK_TIPREL</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Type of the ownership of the savings account in which the loan is deposited (values between 1 and 48). 1 means it is an account holder. Other values mean other type of relationships (i.e. "authorized person but not an owner of the account").</strong></p> </td> </tr> <tr> <td> <p><strong>FK_ORDREL</strong></p> </td> <td> <p><strong>1</strong></p> </td> <td> <p><strong>Order of the ownership relationship. If there are more than account holder, in which position is the FK_NUMPERSO.</strong></p> </td> </tr> </tbody> </table> <p> </p>
UDP Synthetic Dataset for training ML time series models
<p>The dataset available has been produced by the "Next-Generation IoT solutions for the universal supply chain" (iNGENIOUS) project’s consortium under EC grant agreement 957216, made publicly available as part of the Horizon 2020 Open Research Data Pilot (<a href="https://www.openaire.eu/what-is-the-open-research-data-pilot">ORD pilot</a>).<br> The European Commission is not liable for any use that may be made of the information contained herein.</p> <p>The available dataset is in csv format and contains synthetic data of UDP packets received and sent by a single User Plane Function (UPF) covering a span of 6 weeks. The format of the datafile is:</p> <ul> <li>index</li> <li>timestamp </li> <li>UDP packets_rcvd - Total number of UDP packets received</li> <li>UDP packets sent - Total number of UDP packets sent</li> </ul> <p>The simulation was performed based on behavior of UPF and 5GC Network functions inferred from stress tests performed in the iNGENIOUS project's Automated Robots with Heterogeneous Networks Use Case, as well as patterns in urban mobility taken from available UE datasets [NCS+19].</p> <p>More information on the iNGENIOUS project can be found on the project’s website: <a href="https://ingenious-iot.eu/">https://ingenious-iot.eu/</a></p> <p>[NCS+19] Noussan M, Carioni G, Sanvito FD, Colombo E. Urban Mobility Demand Profiles:<br> Time Series for Cars and Bike-Sharing Use as a Resource for Transport and Energy<br> Modeling. Data. 2019; 4(3):108. https://doi.org/10.3390/data4030108</p>
Synthetic Particle Image Dataset (SPID)
<p>SPID is a comprehensive dataset composed of synthetic particle image velocimetry (PIV) image pairs and their corresponding exact optical flow computations. It serves as a valuable resource for researchers and practitioners in the field. The dataset is organized into three subsets: training, validation, and test, distributed in a ratio of 70%, 15%, and 15%, respectively.</p><p>Each subset within SPID consists of an input denoted as "x", which comprises synthetic image pairs. These image pairs provide the necessary context for the optical flow computations. Additionally, an output termed "y" is provided, which represents the exact optical flow calculated for each image pair. Notably, the images within the dataset are single-channel, and the optical flow is decomposed into its u and v components.</p><p>The shape of the input subsets in SPID is given by (number of samples, number of frames, image width, image height, number of channels), representing the dimensions of the input data. On the other hand, the shape of the output subsets is given by (number of samples, velocity components, image width, image height), denoting the shape of the optical flow data.</p><p>It is important to mention that SPID dataset is a preprocessed version of the Raw Synthetic Particle Image Dataset (RSPID), ensuring improved usability and reliability. Moreover, the dataset is packaged as a NumPy compressed NPZ file, which conveniently stores the inputs and outputs as separate NumPy NPZ files with the labels train, validation and test as acess keys. This format simplifies data extraction and integration into machine learning frameworks and libraries, facilitating seamless usage of the dataset.</p><p>SPID incorporates various factors that impact PIV analysis to provide a comprehensive and realistic simulation. The dataset includes image pairs with an image width of 665 pixels and an image height of 630 pixels, ensuring a high level of detail and accuracy with an 8-bit depth. It incorporates different particle radii (1, 2, 3, and 4 pixels) and particle densities (15, 17, 20, 23, 25, and 32 particles) to capture diverse particle configurations.</p><p>To simulate real-world scenarios, SPID introduces displacement variations through the delta x factor, ranging from 0.05% to 0.25%. Noise levels (1, 5, 10, and 15) are also incorporated to mimic practical PIV measurements with varying degrees of noise. Furthermore, out-of-plane motion effects are considered with standard deviations of 0.01, 0.025, and 0.05 to assess their impact on optical flow accuracy.</p><p>The dataset covers a wide range of flow patterns encountered in fluid dynamics. It includes Rankine uniform, Rankine vortex, parabolic, stagnation, shear, and decaying vortex flows, allowing for comprehensive testing and evaluation of PIV algorithms across different scenarios.</p><p>By leveraging the SPID dataset, researchers can develop and validate PIV algorithms and techniques under various challenging conditions. Its realistic and diverse simulation of particle image velocimetry scenarios makes it an invaluable tool for advancing the field and improving the accuracy and reliability of optical flow computations.</p><p> </p>
Synthetic Multimodal Dataset for Daily Life Activities
<p><strong>Outline</strong></p> <ul> <li>This dataset is originally created for the <a href="https://challenge.knowledge-graph.jp/2022/">Knowledge Graph Reasoning Challenge for Social Issue</a>s (KGRC4SI)</li> <li>Video data that simulates daily life actions in a virtual space from Scenario Data.</li> <li>Knowledge graphs, and transcriptions of the Video Data content ("who" did what "action" with what "object," when and where, and the resulting "state" or "position" of the object).</li> <li>Knowledge Graph Embedding Data are created for reasoning based on machine learning </li> <li>This data is open to the public as open data</li> </ul> <p><strong>Details</strong></p> <ul> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Movie">Videos</a></p> <ul> <li>mp4 format</li> <li>203 action scenarios</li> <li>For each scenario, there is a character rear view (file name ending in 0), an indoor camera switching view (file name ending in 1), and a fixed camera view placed in each corner of the room (file name ending in 2-5). Also, for each action scenario, data was generated for a minimum of 1 to a maximum of 7 patterns with different room layouts (scenes). A total of 1,218 videos</li> <li>Videos with slowly moving characters simulate the movements of elderly people.</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/RDF">Knowledge Graphs</a></p> <ul> <li>RDF format</li> <li>203 knowledge graphs corresponding to the videos</li> <li>Includes schema and location supplement information</li> <li>The schema is described below</li> <li><a href="http://kgrc4si.ml:7200/sparql">SPARQL endpoints</a> and <a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/tree/kgrc4si#%E3%83%8A%E3%83%AC%E3%83%83%E3%82%B8%E3%82%B0%E3%83%A9%E3%83%95%E3%81%AE%E4%BD%BF%E7%94%A8%E6%96%B9%E6%B3%95">query examples</a> are available</li> </ul> </li> <li> <p><a href="https://github.com/KnowledgeGraphJapan/KGRC-RDF/blob/kgrc4si/Program">Script Data</a></p> <ul> <li>txt format</li> <li>Data provided to VirtualHome2KG to generate videos and knowledge graphs</li> <li>Includes the action title and a brief description in text format.</li> </ul> </li> <li>Embedding <ul> <li>Embedding Vectors in TransE, ComplEx, and RotatE. Created with DGL-KE (<a href="https://dglke.dgl.ai/doc/">https://dglke.dgl.ai/doc/</a>)</li> <li>Embedding Vectors created with jRDF2vec (<a href="https://github.com/dwslab/jRDF2Vec">https://github.com/dwslab/jRDF2Vec</a>).</li> </ul> </li> </ul> <p><strong>Specification of Ontology</strong></p> <ul> <li>Please refer to the specification for descriptions of all classes, instances, and properties: <a href="https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.html">https://aistairc.github.io/VirtualHome2KG/vh2kg_ontology.htm</a></li> </ul> <p><strong>Related Resources</strong></p> <ul> <li><a href="https://www.youtube.com/watch?v=Ajbn8hNXiZ8&list=PLHaRK-B0LUwjvrPgmIBTrf3DsPhmdnFTW">KGRC4SI Final Presentations with automatic English subtitles (YouTube)</a></li> <li><a href="https://github.com/aistairc/VirtualHome2KG">VirtualHome2KG (Software)</a></li> <li><a href="https://github.com/aistairc/virtualhome_unity_aist">VirtualHome-AIST (Unity</a>)</li> <li><a href="https://github.com/aistairc/virtualhome_aist">VirtualHome-AIST (Python API</a>)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_visualization">Visualization Tool</a> (Software)</li> <li><a href="https://github.com/aistairc/virtualhome2kg_generation">Script Editor</a> (Software)</li> </ul>
Dataset of reported synthetic conditions for ZIF-8 since 2006 and final characteristics and properties of the particles obtained
<p>This is a dataset generated by mining available literature between 2006 and August 2021. Initial literature search was done using Scopus database, employing a combination of keywords such: "ZIF-8", "Zeolitic+ZIF-8", "Framework+ZIF-8", "Zeolitic+MOF".</p> <p>First, removal of duplicates and review articles was done and the remaining documents were chosen by title+abstract analysis. The dataset contains a total of 254 entries, i.e., 254 individual reported synthesis.</p> <p>Throughout this dataset, the following parameters can be found:</p> <p>-About synthesis conditions: Zinc source and the amount employed for the synthesis in mmol (milli-moles); 2-methylimidazole in mmol (HmIm); solvent and quantity employed (in mmol). Modulator and quantity employed (in mmol). Please note that quantities in milli-moles were calculated by hand in most of the cases, since reported data was expressed in different units. Reaction temperature (in °C), reaction time (min) and stirring condition (YES-NO-Initial-time). Finally, reports were classified as "systematic" (or not) based on wether the scope of the work was to explore different synthetic conditions.</p> <p>-About ZIF-8 characteristics: information is mainly focusing on structure-related characteristic, namely the particle morphology (classified as Faceted, poor-faceted, quasispherical, aggregated) and the particle´s size. Reported sizes where classified by the different techniques employed. Finally, Surface area determined by BET formalism was also included.</p>
3D-rhi-synth-2000- Synthetic Rhinophyma Visual Dataset
<p>In the real world, only a handful of data is available for the Rhinophyma skin condition, typically numbering in the hundreds. This repository contains a Synthetic Dataset of Rhinophyma, generated through 3D head models of one male and one female. The purpose of this data generation is to address the data scarcity of the Rhinophyma skin condition within the medical visual data and computer vision community. By generating such data, we aim to bridge the gap in data scarcity for this disease condition, as well as introduce a proof-of-concept methodology for generating synthetic data for specialized disease conditions.</p> <p>The <code>highlight</code> folder 'highlight_female_male_rendered' provides a glimpse of the entire dataset. The file <code>'2000_deformations.npy</code>' contains the 2000 values of deformations applied during rendering.</p> <p>The dataset is divided into two main folders: 'female_rendered' and 'male_rendered'. Within each of these folders, there are three subfolders: 'configu', 'images', and 'points.</p> <p>1. 'configu': This subfolder contains `.json` files with configuration details for each model. The files include various parameters, such as:<br> - "total_num_cameras": the total number of cameras.<br> - "active_camera_name": the name of the active camera.<br> - "camera_focal_len": the camera's focal length.<br> - "camera_loc": the camera's location.<br> - "camera_rot": the camera's rotation.<br> - "nose_deformation_severity": a measure of the severity of nose deformation.<br> - "label": the label for the model (e.g., "Severe").<br> - "nose_variants": additional details about nose variants.</p> <p>2. 'images': This subfolder contains the rendered images in resolution 960x540. They are named according to the following convention e.g.'Nose_Deformation_Severity_0_2.716669764843742_Camera_00001', with specific details related to the deformation severity and camera number. There are images for 10 different cameras.</p> <p>3. 'points': This subfolder contains polygon files corresponding to each model. These files represent the deformations applied to the models during rendering.</p> <p>|-- Dataset Root<br> |-- female_rendered<br> | |-- configu<br> | | |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00001.json<br> | | |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00002.json<br> | | |-- ...<br> | |<br> | |-- images<br> | | |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00001.png<br> | | |-- Nose_Deformation_Severity_0_2.212930927821943_Camera_00002.png<br> | | |-- ...<br> | |<br> | |-- points<br> | | |-- Nose_Deformation_Severity_0_2.212930927821943.ply<br> | | |-- ...<br> | <br> |-- male_rendered<br> |-- configu<br> | |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00001.json<br> | |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00002.json<br> | |-- ...<br> |<br> |-- images<br> | |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00001.jpg<br> | |-- Nose_Deformation_Severity_0_2.716669764843742_Camera_00002.jpg<br> | |-- ...<br> |<br> |-- points<br> |-- Nose_Deformation_Severity_0_2.716669764843742.ply<br> |-- ...</p> <p>Together, these folders and files comprise a dataset designed to represent and analyze the Rhinophyma condition in both male and female 3D head models that we have created. These 3D models will be made available upon a genuine request to the authors of this dataset.</p>
Synthetic AIS Dataset of Vessel Proximity Events
<p>The Automatic Identification System (AIS) allows vessels to share identification, characteristics, and location data through self-reporting. This information is periodically broadcast and can be received by other vessels with AIS transceivers, as well as ground or satellite sensors. Since the International Maritime Organisation (IMO) mandated AIS for vessels above 300 gross tonnage, extensive datasets have emerged, becoming a valuable resource for maritime intelligence.</p> <p>Maritime collisions occur when two vessels collide or when a vessel collides with a floating or stationary object, such as an iceberg. Maritime collisions hold significant importance in the realm of marine accidents for several reasons:</p> <ol> <li>Injuries and fatalities of vessel crew members and passengers.</li> <li>Environmental effects, especially in cases involving large tanker ships and oil spills.</li> <li>Direct and indirect economic losses on local communities near the accident area.</li> <li>Adverse financial consequences for ship owners, insurance companies and cargo owners including vessel loss and penalties.</li> </ol> <p>As sea routes become more congested and vessel speeds increase, the likelihood of significant accidents during a ship's operational life rises. The increasing congestion on sea lanes elevates the probability of accidents and especially collisions between vessels.</p> <p>The development of solutions and models for the analysis, early detection and mitigation of vessel collision events is a significant step towards ensuring future maritime safety. In this context, a synthetic vessel proximity event dataset is created using real vessel AIS messages. The synthetic dataset of trajectories with reconstructed timestamps is generated so that a pair of trajectories reach simultaneously their intersection point, simulating an unintended proximity event (collision close call). The dataset aims to provide a basis for the development of methods for the detection and mitigation of maritime collisions and proximity events, as well as the study and training of vessel crews in simulator environments.</p> <p>The dataset consists of 4658 samples/AIS messages of 213 unique vessels from the Aegean Sea. The steps that were followed to create the collision dataset are:</p> <p>Given 2 vessels X (vessel_id1) and Y (vessel_id2) with their current known location (LATITUDE [lat], LONGITUDE [lon]): </p> <ol> <li>Check if the trajectories of vessels X and Y are spatially intersecting.</li> <li>If the trajectories of vessels X and Y are intersecting, then align temporally the timestamp of vessel Y at the intersect point according to X’s timestamp at the intersect point. The temporal alignment is performed so the spatial intersection (nearest proximity point) occurs at the same time for both vessels.</li> <li>Also for each vessel pair the timestamp of the proximity event is different from a proximity event that occurs later so that different vessel trajectory pairs do not overlap temporarily.</li> </ol> <p>Two csv files are provided. vessel_positions.csv includes the AIS positions vessel_id, t, lon, lat, heading, course, speed of all vessels. Simulated_vessel_proximity_events.csv includes the id, position and timestamp of each identified proximity event along with the vessel_id number of the associated vessels. The final sum of unintended proximity events in the dataset is 237. Examples of unintended vessel proximity events are visualized in the respective png and gif files.</p> <p>The research leading to these results has received funding from the European Union's Horizon Europe Programme under the CREXDATA Project, grant agreement n° 101092749. </p>
Synthetic magnetic nanoparticles for remote-controlled stemcell therapies of neurodegenerative disorders
<p>In the context of the MAGNEURON european project, we developed different types of magnetic nanoparticles that can act as nanoactuators to manipulate intracellular proteins involved in signaling pathways.</p> <p>Four types of particles are presented here. First, size-sorted maghemite cores of different diameter (8 to 20 nm) were synthesized. Then these cores were used to make Fe2O3@SiO2 core-shell nanoparticles that are colloidally stable and easy to functionalize, and poly(acrylic acid) coated nanoparticles. Both types of particles can be rendered fluorescent by the addition of a fluorophore. Finally, we also developed a way to synthesize micro-needles made of aligned maghemite cores encapsulated in a silica layer.</p> <p>In this dataset are presented some electron microscopy images of the optimized particles and their characterizations in terms of sizes and magnetic properties. These particles have then been used by the other members of the Magneuron consortium in order to manipulate different intracellular signalling pathways.</p>
Synthetic spherical tissue images
<p>A time-lapse sequence of synthetic cell membrane images along with their segmentations. Cell-lineages associating cell labels from consecutive time points are also provided.</p> <p> </p> <p><strong>File information:</strong></p> <p>Image files are in the .inr.gz format, which can be read for instance using the <strong>timagetk</strong> (<a href="https://gitlab.inria.fr/mosaic/timagetk">https://gitlab.inria.fr/mosaic/timagetk</a>) Python library. </p>
I-BiDaaS - TID - Synthetic Mobility Data
<p>This is a synthetic data stream based on real-time, cell network events. These events are picked up by the antennas that are closer to the mobile phone thus providing an approximate location of the device. Every transaction of a mobile phone generates one of those events. A transaction can be, for instance, placing or receiving a call, sending or receiving an SMS, asking for a specific URL in your mobile phone browser, or sending a text message or a data transaction from/to any mobile phone app. There are also some synchronization events like, for instance, turning your mobile phone on or off, or when switching between location area networks (relatively big geographical areas comprising several cell towers).</p>
Synthetic data starch potato system Veenkoloniën
<p>Synthetic generated with a simulation model for the starch potato production systems in the Veenkoloniën</p>
Synthetic data for the Bourbonnais
<p>Synthetic data produced using a system dynamics model for beef production systems in the Bourbonnais France.</p>
Synthetically Spoken COCO
<p>Synthetically Spoken COCO</p> <p>Version 1.0</p> <p>This dataset contain synthetically generated spoken versions of MS COCO [1] captions. This<br> dataset was created as part the research reported in [5].<br> The speech was generated using gTTS [2]. The dataset consists of the following files:</p> <p>- dataset.json: Captions associated with MS COCO images. This information comes from [3]. <br> - sentid.txt: List of caption IDs. This file can be used to locate MFCC features of the MP3 files<br> in the numpy array stored in dataset.mfcc.npy.<br> - mp3.tgz: MP3 files with the audio. Each file name corresponds to caption ID in dataset.json<br> and in sentid.txt.<br> - dataset.mfcc.npy: Numpy array with the Mel Frequence Cepstral Coefficients extracted from <br> the audio. Each row corresponds to a caption. The order or the captions corresponds to the<br> ordering in the file sentid.txt. MFCCs were extracted using [4].</p> <p>[1] http://mscoco.org/dataset/#overview<br> [2] https://pypi.python.org/pypi/gTTS<br> [3] https://github.com/karpathy/neuraltalk<br> [4] https://github.com/jameslyons/python_speech_features<br> [5] https://arxiv.org/abs/1702.01991</p>
Synthetic mobile service traffic time series
<p>This dataset contains synthetic mobile service traffic time series used in the paper titled "kaNSaaS: Combining Deep Learning and Optimization for Practical Overbooking of Network Slices", presented at ACM MobiHoc 2023 in Washington, USA. It is composed of 20 time series representing the fluctuations of demands for diverse services categorized under 5G types, including enhanced Mobile Broadband (eMBB), ultra-Reliable Low Latency Communication (uRLLC), and massive Machine Type Communication (mMTC). The time series cover a period of XXX days, and were shown to yield similar properties as those observed in real-world traffic collected in a production mobile network.</p>
Unraveling a synthetic rescue process involved in persisters formation
<p>A common strategy that bacteria utilize to increase their survival under stressful conditions in their natural environments, including antibiotic treatment, is the entry into quiescence, a state of reversible cell growth arrest that offers protection against many environmental insults. Understanding quiescence is an important fundamental question, with relevance in the medical and environmental fields. Little is known about the molecular and physiological determinants that orchestrate survival during this temporary arrest of proliferation, or those that allow a rapid transition back to the proliferating state when conditions again become favorable. In the wide host-range pathogen <em>Salmonella enterica</em> serovar Typhimurium (<em>S</em>. Typhimurium) and other Gram-negative bacteria, this temporary arrest of proliferation induces the expression of the alternative sigma subunit of RNA polymerase, sigma S/RpoS, which remodels global gene expression to reshape the cell physiology and ensure survival under starvation and various stress conditions (<em>i.e. </em>the general stress response).</p> <p>In <em>S</em>. Typhimurium, sigma S is required for stress resistance, biofilm formation and virulence. The incidence of human infections by non-typhoidal <em>Salmonella </em>such as <em>S.</em> Typhimurium increases. These serovars can infect farm animals, thus contaminating animal products, and can be transferred from animal carriers to the environment through fecal matter where they contaminate vegetables, fruits, nuts, and roots. Foodborne diseases caused by <em>Salmonella </em>represent a severe problem to the food supply as well as the public health. <em>S</em>. Typhimurium actively cycles through host and nonhost environments, where it is exposed to a wide variety of stresses, and where sigma S likely plays a crucial role in its persistence.</p> <p>One important aspect of persistence is the phenotypic differentiation of quiescent populations into sub-population(s) of "persisters" that survive in the presence of lethal concentrations of antibiotics. This phenomenon is worsening the worldwide antibiotic crisis, by causing therapy failure and chronic infections and potentially favoring the development of antibiotic resistance. Understanding mechanisms governing bacterial persisters is thus an important topic and a key issue for drug developments. However, despites many studies, the physiological and molecular mechanisms controlling the formation of persisters are poorly understood and controversial. <strong>In the present study we explore the recently discovered and unexpected functional interaction between sigma S and succinate dehydrogenase (Sdh), in the formation of persisters. </strong></p> <p>Persisters are phenotypic variants within a population that survive in the presence of lethal concentrations of antibiotics. When a bacterial population is diluted into fresh medium containing bactericidal antibiotics, a biphasic killing is observed. The bulk of the population, consisting of sensitive cells, dies rapidly and the surviving persisters are killed much more slowly, or do not die during the time course of the experiment. Their regrowth after antibiotic removal yields a new population that has the same sensibility to the antibiotics, as did the parental one. Persisters formation is critically dependent on the growth phase. Despite the identification of a number of genes and pathways involved in persisters formation (toxin-antitoxin modules, SOS and stringent responses, efflux systems, metabolic functions and global regulators), the underlying molecular mechanisms are still poorly understood and controversial. In particular, persisters formation in <em>Escherichia coli</em> K-12 has been reported to be increased, decreased or not affected by a <em>rpoS</em> deletion. A better understanding of physiological parameters favoring persisters formation is critical to develop antipersisters strategies.</p> <p>Succinate dehydrogenase (Sdh), a membrane bound complex that connects the TCA cycle and respiratory chain, is one major target down regulated by sigma S. Negative regulation by sigma S likely targets housekeeping genes that might be deleterious when fully expressed in quiescent cells. Understanding why full expression of those genes has a fitness cost might provide insights into survival mechanisms and weaknesses of quiescent cells, with potential application for antibacterial strategies. To tackle this issue, we used Sdh as a model system. <strong>Our study led us to unravel a synthetic rescue process of a </strong><em><strong>sdh</strong></em><strong> mutation in the formation of persisters. </strong></p> <p>Stationary-phase <em>Salmonella </em>form persisters with a higher frequency than actively growing bacteria, after transfer to fresh medium in the presence of lethal concentrations of ampicillin and ciprofloxacin, but not significant effect of the <em>rpoS</em> deletion on this phenomenon was observed. Surprisingly however, the <em>rpoS</em> deletion suppressed the defect in persister formation of a <em>sdh</em> deletion mutant. Similar results were obtained with independent <em>sdh</em> and <em>sdhrpoS</em> constructs and the mutations did not significantly affect the minimum inhibitory concentration (MIC) of <em>Salmonella</em> for ampicillin and ciprofloxacin. <strong>It is very likely that the </strong><em><strong>rpoS</strong></em><strong> deletion compensates for a metabolic perturbation provoked by the </strong><em><strong>sdh</strong></em><strong> deletion, and key for persister formation. </strong>Since a <em>sdh</em> mutation also decreases persister formation by <em>E. coli </em>and <em>Staphylococcus aureus</em>, the underlying physiological effect might be common to Gram-negative and Gram-positive bacteria. Understanding the molecular and physiological bases of this phenomenon should provide insights into key features driving persisters formation and revival. </p> <p>For references , see the <strong>STUDY</strong> pdf file. </p> <p>The persister assay is described in the <strong>PROTOCOL</strong> pdf file.</p> <p>Strains used are described in the <strong>STRAINS AND PRIMERS</strong> .xlsx file.</p> <p>Results are summarized in the <strong>STUDY</strong> pdf file </p> <p>For detailed data, see the <strong>M1 to M101 persisters</strong> .xlsx files. </p> <p><strong>This work was supported by the French National Research Agency (ANR-19-CE44-0005-01, PERIOMET project).</strong></p> <p>See also:</p> <p>NOREL Francoise, MONTEIL Veronique, DOUCHE Thibaut, & MATONDO Mariette. (2023). Global effects of deletion of the <em>sdh</em> genes, encoding succinate dehydrogenase, and of cobalt on protein abundance in stationary phase <em>Salmonella enterica</em> serovar Typhimurium. [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.8279681">https://doi.org/10.5281/zenodo.8279681</a>. </p> <p>Phégnon, L., Uttenweiler-Joseph, S., & Létisse, F. (2024). <span>Key physiological and metabolic characteristics for the differentiation of quiescent Salmonella's cells into persisters [Data set]. </span>Zenodo. <a href="https://doi.org/10.5281/zenodo.10885905" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.10885905</a></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.