Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
250
datasets available to search
ShareScore release 0.9.0
Dataset results
250 results for “Synthetic dataset”
Datasets of synthetic workflows for evaluating a multi-objective and multi-constrained scheduling approach for cyber-physical applications
<p>These datasets of synthetic workflows (task graphs) were generated to evaluate the performance and scalability of a multi-objective and multi-constrained scheduling approach for workflow applications of various structures, sizes, and sensing/actuating requirements in a cyber-physical system (CPS) based on the edge-hub-cloud paradigm. The examined CPS comprised four edge devices (i.e., single-board computers, each attached to an unmanned aerial vehicle (UAV) equipped with sensors/actuators) interacting with a hub device (e.g., a laptop), which in turn communicated with a more computationally capable cloud server. All system devices featured heterogeneous multicore processors with different processing core failure rates and varied sensing/actuating or other specialized capabilities. Our objectives were the minimization of the overall latency, the minimization of the overall energy consumption, and the maximization of the overall reliability of the workflow application in the specific CPS, under deadline, reliability, memory, storage, energy, capability, and task precedence constraints.</p> <p>We generated 25 random task graphs with 10, 20, 30, 40, and 50 nodes (5 task graphs for each size), utilizing the Task Graphs For Free (TGFF) random task graph generator [1],[2]. Additional task parameters (e.g., execution time, power consumption, memory, storage, output data size, capability, reliability threshold) were included post-generation, using appropriate values. More details are provided in README.txt.<br><br>References:<br>[1] R. P. Dick, D. L. Rhodes, and W. Wolf, "TGFF: Task graphs for free," Proceedings of the Sixth International Workshop on Hardware/Software Codesign (CODES/CASHE), 1998, pp. 97-101, doi: 10.1109/HSC.1998.666245.<br>[2] R. P. Dick, D. L. Rhodes, and K. Vallerio, "TGFF," https://robertdick.org/projects/tgff/.</p>
Synthetic Multimodal Drone Delivery Dataset
<h1><strong>README: Synthetic Logistics Dataset Structure and Components</strong></h1> <p>This dataset provides a structured representation of logistics data designed to evaluate and optimize hybrid truck-and-drone delivery networks. It captures a comprehensive set of parameters essential for modeling real-world logistics scenarios, including spatial coordinates, environmental conditions, and operational constraints. The data is meticulously organized into distinct keys, each representing a critical aspect of the delivery network, enabling researchers and practitioners to conduct flexible and in-depth analyses. </p> <p>The dataset is a curated subset derived from the research presented in the paper "Synthetic Dataset Generation for Optimizing Multimodal Drone Delivery Systems" by Altinsel et al. (2024), published in Drones. It serves as a practical resource for studying the interplay between ground-based and aerial delivery systems, with a focus on efficiency, environmental impact, and operational feasibility. </p> <p><strong>Altinses, D., Torres, D. O. S., Gobachew, A. M., Lier, S., & Schwung, A. (2024). Synthetic Dataset Generation for Optimizing Multimodal Drone Delivery Systems. <em>Drones (2504-446X)</em>, <em>8</em>(12).</strong></p> <p>Each data file contains information on ten customer locations, specified by their x and y coordinates, which facilitate the modeling of delivery routes and service areas. Additionally, the dataset includes communication data represented as a two-dimensional grid, which can be used to assess signal strength, connectivity, or other network-related factors that influence drone operations. </p> <p>A key feature of this dataset is the inclusion of wind data, structured as a two-dimensional grid with four distinct features per grid point. These features likely represent wind velocity components (such as horizontal and vertical directions) along with auxiliary parameters like turbulence intensity or wind shear, which are crucial for drone path planning and energy consumption estimation. The wind data enables researchers to simulate realistic environmental conditions and evaluate their impact on drone performance, stability, and battery life. </p> <p>By integrating geospatial, environmental, and operational data, this dataset supports a wide range of applications, from route optimization and energy efficiency studies to risk assessment and resilience planning in multimodal delivery systems. Its synthetic nature ensures reproducibility while maintaining relevance to real-world logistics challenges, making it a valuable tool for advancing research in drone-assisted delivery networks. </p> <p> </p> <h3><strong>The 4 wind channels represent:</strong></h3> <ol> <li> <p><strong><code>X</code> and <code>Y</code> (Grid Positions)</strong></p> <ul> <li> <p>These define <strong>where the arrows start</strong> (usually a meshgrid).</p> </li> </ul> </li> <li> <p><strong><code>U</code> and <code>V</code> (Arrow Directions)</strong></p> <ul> <li> <p><code>U</code> = Horizontal component (e.g., gradient in <code>x</code>).</p> </li> <li> <p><code>V</code> = Vertical component (e.g., gradient in <code>y</code>).</p> </li> </ul> </li> </ol> <p> </p> <h3><strong>How to load the files using Python:</strong></h3> <p>data = np.loadtxt('data.txt')</p> <p>#### Just for Wind data:</p> <p>data = data.reshape((4,16,16)) </p>
ASVspoof2019LA-Sim: Augmented Dataset for An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems
<p>This is the dataset we augmented to study the channel effects for anti-spoofing. For more details, please refer to our Interspeech 2021 paper: "An Empirical Study on Channel Effects for Synthetic Voice Spoofing Countermeasure Systems".</p> <p>Proceeding: <a href="https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html">https://www.isca-speech.org/archive/interspeech_2021/zhang21ea_interspeech.html</a></p> <p>Arxiv: <a href="https://arxiv.org/pdf/2104.01320.pdf">https://arxiv.org/pdf/2104.01320.pdf</a></p> <p>Code: <a href="https://github.com/yzyouzhang/Empirical-Channel-CM">https://github.com/yzyouzhang/Empirical-Channel-CM</a></p> <p>Contact: you.zhang@rochester.edu</p> <p><strong>Version 1.0</strong> contains the <strong>training</strong> and the <strong>development</strong> set. We have added the <strong>evaluation</strong> set in <strong>version 1.1 </strong>but deleted the training set due to the size limitation, but you can still access the training set in version 1.0.</p> <p>Please check it out.</p> <p>To extract the files, please use the following commands:</p> <pre><code class="language-bash">cat eval.tar.gz-part* > eval.tar.gz tar -xvzf *.tar.gz</code></pre> <p>After concatenation, to make sure the download is complete, you can check with the following:</p> <pre><code>md5sum *.tar.gz 15dea7d28b126994bb6b159778f706af dev.tar.gz 0615052b34ca6c7f58505eaa8647844f eval.tar.gz 3058dd9d407f3c9ae697acca8c34a6c3 train.tar.gz</code></pre> <p>Thanks.</p>
Zellige example dataset: synthetic image dataset
<p><strong>Phantom 3D image containing three distinct and superimposed synthetic surfaces. </strong></p> <p>It models a typical stack of confocal images of epithelial and non-epithelial structures.The surfaces generated are of two types: “solid” surfaces, presenting a homogeneous signal over the entire surface, or surfaces presenting a signal restricted to a polygonal mesh mimicking the mesh of apical cellular junctions of an epithelium observed at its surface. This dataset contains both the ground-truth height maps and the height maps generated with Zellige. The Zellige parameters used are:</p> <p><span class="math-tex">\(T_{A}=16, T_{otsu}=12, S_{min}=5, \sigma_{xy}=4, \sigma_{z}=2, T_{OSE1}=0.9, R_{1}=5, C_{1}=0.1, T_{OSE2}=0.1, R_{2}=10, C_{2}=0.8.\)</span></p> <p>Nota: the ground-truth height maps can be directly compared to Zellige height maps by subtraction.</p> <p>See the accompanying paper: Extracting multiple surfaces from 3D microscopy images in complex biological tissues with the Zellige software tool. Trébeau <em>et al.</em> 2022: <a href="https://doi.org/10.1101/2022.04.05.485876">https://doi.org/10.1101/2022.04.05.485876</a></p>
HIKARI-2021: Generating Network Intrusion Detection Dataset Based on Real and Encrypted Synthetic Attack Traffic
<p>Available datasets from the paper Generating Encrypted Network Traffic for Intrusion Detection Datasets.</p> <p>To produce the dataset follow the technical detail in <a href="https://github.com/andreysfc/generating-encrypted-network">github</a></p>
Datasets of synthetic workflows for cyber-physical edge-hub-cloud systems
<p>These datasets of synthetic workflows were generated to evaluate the performance and scalability of a multi-constrained scheduling approach for workflow applications of various structures, sizes, and sensing/actuating requirements in a cyber-physical system (CPS) following the edge-hub-cloud paradigm. The examined CPS comprised four edge devices (i.e., single-board computers, each attached to an unmanned aerial vehicle (UAV) equipped with sensors/actuators) interacting with a hub device (e.g., a laptop), which in turn communicated with a more computationally capable cloud server. All system devices featured heterogeneous multicore processors and varied sensing/actuating or other specialized capabilities. The problem objective was the minimization of the overall latency of the application under deadline, memory, storage, energy, capability, and task precedence constraints.</p> <p>We generated 25 random workflows (task graphs) with 10, 20, 30, 40, and 50 nodes (5 task graphs for each size), utilizing the Task Graphs For Free (TGFF) random task graph generator [1],[2]. Additional task parameters (e.g., execution time, power consumption, memory, storage, output data size, capability) were included post-generation, using appropriate values. More details are provided in README.txt and in [3].<br><br>References:<br>[1] R. P. Dick, D. L. Rhodes, and W. Wolf, "TGFF: Task graphs for free," in Proc. Sixth International Workshop on Hardware/Software Codesign (CODES/CASHE), 1998, pp. 97-101, doi: 10.1109/HSC.1998.666245.</p> <p>[2] R. P. Dick, D. L. Rhodes, and K. Vallerio, "TGFF," https://robertdick.org/projects/tgff/.</p> <p>[3] A. Kouloumpris, G. L. Stavrinides, M. K. Michael, and T. Theocharides, “Optimal multi-constrained workflow scheduling for cyber-physical systems in the edge-cloud continuum,” in Proc. 2024 IEEE 48th Annual Computers, Software, and Applications Conference (COMPSAC), Jul. 2024, pp. 483-492, doi: 10.1109/COMPSAC61105.2024.00072.</p>
AstroChat - A Dataset of synthetically generated conversations for LLM supervised fine-tuning in the domain of Space Mission Engineering and Astronautics
<h1>AstroChat Dataset Description</h1> <h2>Purpose and Scope</h2> <p>The AstroChat dataset is a collection of 901 dialogues, synthetically generated, tailored to the specific domain of Astronautics / Space Mission Engineering. This dataset will be frequently updated following feedback from the community. If you would like to contribute, please reach out in the community discussion.</p> <h2>Intended Use</h2> <p>The dataset is intended to be used for supervised fine-tuning of chat LLMs (Large Language Models). Due to its currently limited size, you should use a pre-trained instruct model and ideally augment the AstroChat dataset with other datasets in the area of (Science Technology, Engineering and Math).</p> <h2>DATASET DESCRIPTION</h2> <h3>Access</h3> <ul> <li>Manual download from Hugging face hub: <a href="https://huggingface.co/datasets/patrickfleith/Astro-Ultrachat" rel="nofollow">https://huggingface.co/datasets/patrickfleith/AstroChat</a></li> <li>Or with python:</li> </ul> <pre><code>from datasets import load_dataset dataset = load_dataset("patrickfleith/AstroChat") </code></pre> <h3>Structure</h3> <p>901 generated conversations between a simulated user and AI-assistant (more on the generation method below). Each instance is made of the following field (column):</p> <ul> <li><strong>id</strong>: a unique identifier to refer to this specific conversation. Useeful for traceability purposes, especially for further processing task or merge with other datasets.</li> <li><strong>topic</strong>: a topic within the domain of Astronautics / Space Mission Engineering. This field is useful to filter the dataset by topic, or to create a topic-based split.</li> <li><strong>subtopic</strong>: a subtopic of the topic. For instance in the topic of <code>Propulsion</code>, there are subtopics like <code>Injector Design</code>, <code>Combustion Instability</code>, <code>Electric Propulsion</code>, <code>Chemical Propulsion</code>, etc.</li> <li><strong>persona</strong>: description of the persona used to simulate a user</li> <li><strong>opening_question</strong>: the first question asked by the user to start a conversation with the AI-assistant</li> <li><strong>messages</strong>: the whole conversation messages between the user and the AI assistant in already nicely formatted for rapid use with the transformers library. A list of messages where each message is a dictionary with the following fields: <ul> <li><strong>role</strong>: the role of the speaker, either <code>user</code> or <code>assistant</code></li> <li><strong>content</strong>: the message content. For the assistant, it is the answer to the user's question. For the user, it is the question asked to the assistant.</li> </ul> </li> </ul> <p><strong>Important</strong> See the full list of topics and subtopics covered below.</p> <h3>Metadata</h3> <p>Dataset is version controlled and commits history is available here: <a href="https://huggingface.co/datasets/patrickfleith/Astro-Ultrachat/commits/main" rel="nofollow">https://huggingface.co/datasets/patrickfleith/AstroChat/commits/main</a></p> <h3>Generation Method</h3> <p>We used a method inspired from Ultrachat dataset. Especially, we implemented our own version of Human-Model interaction from <strong>Sector I: Questions about the World</strong> of their paper:</p> <p><em>Ding, N., Chen, Y., Xu, B., Qin, Y., Zheng, Z., Hu, S., ... & Zhou, B. (2023). Enhancing chat language models by scaling high-quality instructional conversations. arXiv preprint arXiv:2305.14233.</em></p> <h4>Step-by-step description</h4> <ul> <li>Defined a set of user persona</li> <li>Defined a set of topics/ disciplines within the domain of Astronautics / Space Mission Engineering</li> <li>For each topics, we defined a set of subtopics to narrow down the conversation to more specific and niche conversations (see below the full list)</li> <li>For each subtopic we generate a set of opening questions that the user could ask to start a conversation (see below the full list)</li> <li>We then distil the knowledge of an strong Chat Model (in our case ChatGPT through then api with <code>gpt-4-turbo</code> model) to generate the answers to the opening questions</li> <li>We simulate follow-up questions from the user to the assistant, and the assistant's answers to these questions which builds up the messages.</li> </ul> <h3>Future work and contributions appreciated</h3> <ul> <li>Distil knowledge from more models (Anthropic, Mixtral, GPT-4o, etc...)</li> <li>Implement more creativity in the opening questions and follow-up questions</li> <li>Filter-out questions and conversations which are too similar</li> <li>Ask topic and subtopic expert to validate the generated conversations to have a sense on how reliable is the overall dataset</li> </ul> <h3>Languages</h3> <p>All instances in the dataset are in english</p> <h3>Size</h3> <p>901 synthetically-generated dialogue</p> <h2>USAGE AND GUIDELINES</h2> <h3>License</h3> <p>AstroChat © 2024 by Patrick Fleith is licensed under Creative Commons Attribution 4.0 International</p> <h4>Restrictions</h4> <p>No restriction. Please provide the correct attribution following the license terms.</p> <h4>Citation</h4> <p><em>Patrick Fleith, AstroChat – A Dataset of synthetically generated conversations for LLM supervised fine-tuning in the domain of Space Mission Engineering and Astronautics, (2024).</em></p> <h4>Update Frequency</h4> <p>Will be updated based on feedbacks. I am also looking for contributors. Help me create more datasets for Space Engineering LLMs :)</p> <h4>Have a feedback or spot an error?</h4> <p>Use the community discussion tab directly on the huggingface AstroChat dataset page.</p> <h4>Contact Information</h4> <p>Reach me here on the community tab or on LinkedIn (Patrick Fleith) with a Note.</p> <h3>Number of conversation per topic category</h3> <pre><code>Space Propulsion Systems 135 Human Spaceflight 50 Entry Descent and Landing (EDL) 45 Mechanisms 45 Planetary Rovers 45 Attitude Determination and Control 45 Telecommunication 41 Space Business 40 Structures 40 Materials 40 Launchers, Launches, Launch Operations 36 Power System 35 Payload S/S and Optics 35 Reliability, Availability, Maintainability, and Safety (RAMS) 35 Space Missions Operations 31 Space Environment 30 Command and Data System 30 Orbital Mechanics 30 Space Law 26 Ground Systems 25 Thermal Control 25 Space Processes 20 Planetary Science and Exploration 17 </code></pre> <h3>Topics and subtopics covered</h3> <p>topic: [ Space Law ]</p> <p>subtopics:</p> <ul> <li>Space Law Basics</li> <li>1998 ISS agreement</li> <li>Outer Sppace Treaty</li> <li>Geostationary Orbit Regulations</li> <li>Space Traffic Management</li> <li>French Space Law</li> </ul> <p>topic: [ Space Business ]</p> <p>subtopics:</p> <ul> <li>New Space</li> <li>Satellite Insurance</li> <li>Financing Space Project (in EU)</li> <li>Commercial Satellite Launch Services</li> <li>Space Tourism</li> <li>Business Models for Space Stations</li> <li>Public-private Partnerships</li> <li>Economic Impact of Space Technologies</li> </ul> <p>topic: [ Space Missions Operations ]</p> <p>subtopics:</p> <ul> <li>Flight control team</li> <li>Flight Dynamics</li> <li>Procedure Preparation and Validation</li> <li>Mission Planning</li> <li>Extravehicular Activities (EVAs)</li> <li>Collision Avoidance Manoeuvres</li> <li>Mission Termination and De-Orbit Strategies</li> </ul> <p>topic: [ Human Spaceflight ]</p> <p>subtopics:</p> <ul> <li>Astronaut Selection</li> <li>Astronaut Training</li> <li>research experiments onboard of the ISS</li> <li>Human Mission to Mars Design</li> <li>Environmental Control and Life Support Systems</li> <li>Moon Surface Habitats</li> <li>Microgravity effects</li> <li>Space Suit Design and Operation</li> <li>Space Medicine</li> <li>Space Food</li> </ul> <p>topic: [ Space Environment ]</p> <p>subtopics:</p> <ul> <li>Micrometeorites</li> <li>Space Radiation</li> <li>Solar Cycle</li> <li>Spacecraft Hardening</li> <li>Space Environment Effects on Satellites</li> <li>Magneto-sphere and Radiation Belt</li> </ul> <p>topic: [ Space Propulsion Systems ]</p> <p>subtopics:</p> <ul> <li>Liquid Rocket Engines</li> <li>Solid Rocket Motors</li> <li>Hybrid Rocket Engines</li> <li>Staging and Ignition Systems</li> <li>Propellant Feed Systems</li> <li>Nozzle Designs</li> <li>Thermodynamics</li> <li>Turbopumps and/or Combustion Chambers</li> <li>Specific Impulse and Thrust-to-Weight Ratios</li> <li>Chemical Monopropellant Technologies</li> <li>Chemical Bipropellant Systems</li> <li>Nuclear Thermal Propulsion</li> <li>Fuel Handling and Storage</li> <li>Nuclear Propulsion Thermal Neutron Absorbers</li> <li>Nuclear Propulsion Heat Exchangers</li> <li>Green Propellants</li> <li>Bipropellant Injector Design</li> <li>Electric Ion Thrusters</li> <li>Hall Effect Thrusters</li> <li>Electrothermal Thrusters</li> <li>Grid and Cathode Technologies</li> <li>Aerospike Engines</li> <li>Variable Specific Impulse Magnetoplasma Rocket (VASIMR)</li> <li>Bipropellant Mixing Ratios and Combustion</li> <li>Cryogenic Propellant Handling</li> <li>Oxydizer and Fuel Combinations</li> <li>Long-term Impacts of Propellant Residues in the Atmosphere</li> <li>Propellant Tank Pressurization</li> </ul> <p>topic: [ Space Processes ]</p> <p>subtopics:</p> <ul> <li>Trade Studies</li> <li>Margins, Coningencies, Reserves</li> <li>Systems Engineering</li> <li>Quality Assurance</li> </ul> <p>topic: [ Ground Systems ]</p> <p>subtopics:</p> <ul> <li>Ground Stations</li> <li>Ground Support Equipments</li> <li>Control Centers</li> <li>Tracking Systems</li> <li>AntennasGround Systems Engineering</li> </ul> <p>topic: [ Planetary Rovers ]</p> <p>subtopics:</p> <ul> <li>Mars Rovers</li> <li>Lunar Rovers</li> <li>Rover Instrumentation</li> <li>Rover Power Systems</li> <li>Rover Thermal Control</li> <li>Rover Autonomy</li> <li>Wheels Design</li> <li>Legged Rovers</li> <li>Hazard Avoidance</li> </ul> <p>topic: [ Planetary Science and Exploration ]</p> <p>subtopics:</p> <ul> <li>Astrobiology</li> <li>Exoplanets</li> <li>AsteroidsJupiter</li> <li>Saturn</li> <li>Search for Extraterrestrial Life</li> </ul> <p>topic: [ Structures ]</p> <p>subtopics:</p> <ul> <li>Structural Design and Analysis</li> <li>Load Path Determination</li> <li>Vibration and Acoustic Testing</li> <li>Thermal Protection Systems</li> <li>Composite Structures</li> <li>Joining Techniques (e.g., Welding, Bolting, Bonding)</li> <li>Manufacturing Tolerances and Quality Control</li> <li>Deployable Structures (e.g., Antennas, Solar Arrays)</li> </ul> <p>topic: [ Mechanisms ]</p> <p>subtopics:</p> <ul> <li>Actuators and Dampers</li> <li>Gimbals and Bearings</li> <li>Latch and Release Devices</li> <li>Hinges and Deployment Systems</li> <li>Robotic Arms and Tools</li> <li>Valves and Fluid Control Systems</li> <li>Thermal Expansion Joints</li> <li>Drive Systems and Motors</li> <li>Reliability and Lifetime Analysis</li> </ul> <p>topic: [ Materials ]</p> <p>subtopics:</p> <ul> <li>Composite Materials</li> <li>Metals and Alloys</li> <li>Polymers and Plastics</li> <li>Nano-materials</li> <li>Radiation Shielding Materials</li> <li>Thermal Insulation Materials</li> <li>Corrosion and Oxidation Resistance</li> <li>Material Testing and Characterization</li> </ul> <p>topic: [ Entry Descent and Landing (EDL) ]</p> <p>subtopics:</p> <ul> <li>Aerodynamics and Aeroheating</li> <li>Powered Descent</li> <li>Landing Gear and Systems</li> <li>Heat Shield Design and Materials</li> <li>Hazard Avoidance</li> <li>Surface Interaction (Airbags, Crushable Structures)</li> <li>Entry, Descent, and Landing Sequencing</li> <li>EDL on Mars</li> <li>Parachute Systems Design</li> </ul> <p>topic: [ Reliability, Availability, Maintainability, and Safety (RAMS) ]</p> <p>subtopics:</p> <ul> <li>System Reliability Modeling</li> <li>Failure Modes, Effects, and Criticality Analysis (FMECA)</li> <li>Risk Assessment and Management</li> <li>Safety-Critical Systems Design</li> <li>Availability Modeling and Prediction</li> <li>Lifecycle Cost and Duration Analysis</li> <li>Hazardous Material Handling</li> </ul> <p>topic: [ Orbital Mechanics ]</p> <p>subtopics:</p> <ul> <li>Interplanetary Trajectories</li> <li>Gravity Assist Maneuvers</li> <li>Orbit Determination and Propagation</li> <li>Space Situational Awareness and Debris Tracking</li> <li>Mission Design and Analysis Tools</li> <li>Orbit Decay and Re-entry Predictions</li> </ul> <p>topic: [ Launchers, Launches, Launch Operations ]</p> <p>subtopics:</p> <ul> <li>Launcher Types (e.g., expendable, reusable)</li> <li>Launch Vehicles</li> <li>Launch Sites and Infrastructure</li> <li>Countdown Procedures and Sequencing</li> <li>Launch Window Determination and Trajectory Analysis</li> <li>Ground and Launch Crew Training</li> <li>Payload Integration and Fairing Design</li> <li>Environmental and Weather Constraints</li> </ul> <p>topic: [ Attitude Determination and Control ]</p> <p>subtopics:</p> <ul> <li>Sensors for Attitude Determination (e.g., Gyroscopes, Star Trackers)</li> <li>Actuators for Attitude Control (e.g., Reaction Wheels, Thrusters)</li> <li>Control Algorithms (e.g., PID, Kalman Filter)</li> <li>Momentum Exchange Devices</li> <li>Attitude Dynamics Modeling</li> <li>On-Orbit Attitude Reconfiguration</li> <li>Fault Detection and Response Strategies</li> <li>Sun and Earth Sensors</li> <li>Magnetic Torquers and Gravity Gradient Stabilization</li> </ul> <p>topic: [ Payload S/S and Optics ]</p> <p>subtopics:</p> <ul> <li>Payload Design and Integration</li> <li>Spectral Imaging and Multi-spectral Sensors</li> <li>Infrared and Ultraviolet Optics</li> <li>Calibration and Validation of Optical Systems</li> <li>Image Processing and Data Analysis</li> <li>Thermal Control for Sensitive Optics</li> <li>Data Downlink and Communication Interfaces</li> </ul> <p>topic: [ Power System ]</p> <p>subtopics:</p> <ul> <li>Solar Panels and Arrays</li> <li>Battery Types and Management Systems (e.g., Li-ion, NiMH)</li> <li>Energy Storage Technologies</li> <li>Fault Protection and Isolation</li> <li>Harness and Cabling</li> <li>Alternative Power Sources (e.g., RTGs, Fuel Cells)</li> <li>Power Budgeting and Load Analysis</li> </ul> <p>topic: [ Thermal Control ]</p> <p>subtopics:</p> <ul> <li>Active Thermal Control Systems (e.g., Heat Pumps, Louvers)</li> <li>Environmental Testing and Validation</li> <li>Heating and Cooling Hardware</li> <li>Thermal Protection for Entry, Descent, and Landing</li> <li>Cryogenic Thermal Management</li> </ul> <p>topic: [ Command and Data System ]</p> <p>subtopics:</p> <ul> <li>Onboard Computers and Processing Units</li> <li>Software Architecture and Middleware</li> <li>Command Link and Telemetry Systems</li> <li>Interface and Bus Systems (e.g., MIL-STD-1553, SpaceWire)</li> <li>Real-Time Operating Systems (RTOS)</li> <li>Security Measures and Encryption</li> </ul> <p>topic: [ Telecommunication ]</p> <p>subtopics:</p> <ul> <li>Antenna Systems (e.g., Parabolic, Phased Array)</li> <li>Communication Transponders</li> <li>Frequency Bands and Spectrum Management</li> <li>Signal Modulation and Demodulation Techniques</li> <li>Inter-Satellite Links and Data Relays</li> <li>Error Detection and Correction</li> <li>Space Communication Protocols</li> <li>RF and Microwave Components</li> <li>Deep Space Communications</li> </ul>
iHelp public synthetic dataset
<p>The iHelp public synthetic dataset has been synthesized from primary and secondary data collected during the Medical University Plovdiv pilot implementing their cancer program. A detailed description can be found in the "iHelp public synthetic dataset.docx" file.</p>
Synthetic Dataset of Cardiac Microbundles
<p>The synthetic microbundle dataset consists of 60 200x256x256 ".tif" files generated based on textures extracted from real data and warped according to experimentally-informed Finite Element (FE) simulations. More details about generating this dataset can be found on the dedicated <a href="https://github.com/HibaKob/SyntheticMicroBundle">GitHub repository</a>.</p>
Synthetic Dataset of charging processes by electric vehicles at workplace in Germany
<p>The dataset shows eight cluster groups that depict the mobility behavior of electric vehicle users in the employee context. For this purpose, 23.9 million data entries were analyzed, corresponding to 37,238 charging sessions. These data were collected over the year 2023. The 220 charging points were exclusively accessible to employees (private use case). From the data, cluster groups were derived using the Gaussian Mixture Model, and a synthetic dataset was generated through Monte Carlo sampling.</p> <p><span>The dataset consists of 8000 synthetic profiles, offering a robust scientific basis. By retaining the same statistical attributes as the empirical data, the synthetic profiles represent eight different mobility clusters, each containing 1000 entries, including full-time and part-time employees, shift workers, pool vehicle users, and opportunists.</span> Each cluster is represented by the mean parking start hours (arrival time - in decimal hours), mean parking duration (in decimal hours), the average energy recharged, and the average charging duration, each including the cluster-specific standard deviation and median.</p> <p>Further information can be obtained from the upcoming publication: "Synthetic Dataset of charging processes by electric vehicles at workplace in Germany."</p>
UnLoc dataset (Synthetic + Real)
<p>This dataset contains the synthetic and real data used in the article " Virtual Training for a Real Application: Accurate Object-Robot Relative Localization Without Calibration " to train 3 convolutional neural networks (CNNs) in order to perform uncalibrated relative localization of a cuboid block with respect to a robot, and to evaluate them. </p> <p>It consists of a dataset composed of synthetic pictures for training the CNNs and a dataset of real pictures for evaluation. </p> <p>The "synthetic" dataset is composed 3 sub-datasets (each of them composed of thousands of synthetic pictures and corresponding groundtruth) for training : </p> <ol> <li>a dataset for coarse relative localization subtask : coarse_estimation_data (~14.6 GB when extracted)</li> <li>a dataset for the tool localization subtask : tool_detection_data (~4.3 GB when extracted)</li> <li>a dataset for the fine relative localization subtask : fine_estimation_data (~13.3 GB when extracted)</li> </ol> <p>These sub-datasets are composed of raw data as well as post-treated data ready to be used for training CNNs, in CSV format and in Torch format (.t7). </p> <p>The "real" dataset is composed of real pictures with a precisely localized cuboid block for evaluation only : UnLoc_real (~2.5 GB when extracted). </p> <p>More information available on the project page : <a href="http://imagine.enpc.fr/~loingvi/unloc/">http://imagine.enpc.fr/~loingvi/unloc/</a></p>
A Dataset with Synthetic Landing Trajectories for Zurich Airport
<p>The archive contains synthetic datasets in npy format generated using a TimeGAN-based model designed to capture a range of aircraft landing trajectories at Zurich airport across different operational scenarios and environmental conditions. Each dataset consists of multiple groups representing distinct patterns or behaviors in the trajectory data. The trajectories are segmented in clusters, and to some of them a smoothing filter was applied.</p> <p><span>The datasets incorporate a range of variables critical for modeling aircraft landing behaviors. Continuous variables such as longitude, latitude, and altitude exhibit multimodal distributions, capturing different operational phases and conditions within each cluster. The data is stored in an array format with dimensions (number of samples, sequence length, feature dimensions). Here, the number of samples corresponds to the total number of recorded flight trajectories included in the dataset, while the sequence length represents the duration or the number of time steps over which each trajectory is recorded. The feature dimensions denote the various variables (state vector) measured at each time step, consisting of longitude, latitude and altitude. Categorical variables, such as runway identifiers and cluster labels, follow distributions that reflect operational frequencies, with certain clusters or runways being more common under specific conditions. </span></p> <p>The archive contains the following files:</p> <p>- 5clust0.npy, 5clust1.npy, 5clust2.npy, 5clust3.np & 5clust4.npy (5 clusters of landing trajectories separated)<br>- ma_5clust0.npy, ma_5clust1.npy, ma_5clust2.npy, ma_5clust3.np & ma_5clust4.npy (5 clusters of landing trajectories separated, moving average filter applied)<br>- ma_3clust0.npy, ma_3clust1.npy & ma_3clust2.npy (3 clusters of landing trajectories separated, moving average filter applied)<br>- run28_syn.npy & run24_syn.npy (groups of landing trajectories per runway)<br>- ma_run28_syn.npy & ma_run24_syn.npy (groups of landing trajectories per runway, moving average filter applied)<br>- go_around_synthetic.npy (go-around landing trajectories on runway 14)</p>
Synthetic dataset of user interactions - postpartum depression.csv
<p>A synthetic data set composed of 200 users' utterances as possible answers to questions related to these topics:</p> <p>(i) Feeling sad or Tearful<br>(ii) Irritable towards baby & partner<br>(iii) Trouble sleeping at night<br>(iv) Problems concentrating or making decision<br>(v) Overeating or loss of appetite<br>(vi) Feeling of guilt<br>(vii) Problems of bonding with baby <br>(viii) Suicide attempt</p>
A Synthetic Hyperspectral Dataset for Development and Validation of Phytoplankton Size Class Retrieval Models
<p><strong>A Synthetic Hyperspectral Dataset for Development and Validation of Phytoplankton Size Class Retrieval Models.</strong></p> <p>Please refer to the following scientific paper for a description of the dataset.</p> <blockquote> <p>Holtrop, T.; Van Der Woerd, H.J. (accepted) HYDROPT: An Open-Source Framework for Fast Inverse Modelling of Multi- and Hyperspectral Observations from Oceans, Coastal and Inland Waters. <em>Remote Sens. </em><strong>2021</strong>, 13, 0.</p> </blockquote>
Synthetic dataset used for validating MDSPACE method for analyzing continuous conformational variability of biomolecules in cryo-EM single particle images
<p>Synthetic dataset used for validating MDSPACE method for analyzing continuous conformational variability of biomolecules in cryo-EM single particle images. A README file with the contents of the dataset is included. </p>
UDAPDR Synthetic Queries Datasets Sample
<p>Sample of synthetic queries datasets for <a href="https://arxiv.org/abs/2303.00807">UDAPDR: Unsupervised Domain Adaptation via LLM Prompting and Distillation of Rerankers</a>.</p>
Raw Synthetic Particle Image Dataset (RSPID)
<p>Synthetic Particle Image Velocimetry (PIV) data generated by PIV Image Generator Software. Which is a tool that generates synthetic Particle Imaging Velocimetry (PIV) images with the purpose of validating and benchmarking PIV and Optical Flow methods in tracer based imaging for fluid mechanics (Mendes et al., 2020). </p><p>This data was generated with the following parameters:</p><ul><li>image width: 665 pixels;</li><li>image height: 630 pixels;</li><li>bit depth: 8 bits;</li><li>particle radius: 1, 2, 3, 4 pixels;</li><li>particle density: 15, 17, 20, 23, 25, 32 particles;</li><li>delta x factor: 0.05, 0.1, 0.15, 0.2, 0.25 %;</li><li>noise level: 1, 5, 10, 15;</li><li>out-of-plane standard deviation: 0.01, 0.025, 0.05;</li><li>flows: rankine uniform, rankine vortex, parabolic, stagnation, shear, decaying vortex.</li></ul>
CBTex - A Dataset of Synthetic Cardboard Textures
<p>Dataset of >200 synthetic cardboard texture images that were rendered with <a href="https://www.blendermarket.com/creators/doublegum">DoubeGum</a>'s <a href="https://www.blendermarket.com/products/cardboard">cardboard shader</a> in Blender. Used to generate Parcel3D, the dataset for our <a href="https://openaccess.thecvf.com/content/CVPR2023W/VISION/html/Naumann_Parcel3D_Shape_Reconstruction_From_Single_RGB_Images_for_Applications_in_CVPRW_2023_paper.html">paper</a> on single image 3D reconstructions of potentially damaged parcels. See our <a href="https://a-nau.github.io/parcel3d/">project page</a> for details.</p> <p> </p> <p>If you use this resource for scientific research, please consider citing</p> <pre><code>@inproceedings{naumannParcel3DShapeReconstruction2023, author = {Naumann, Alexander and Hertlein, Felix and D\"orr, Laura and Furmans, Kai}, title = {Parcel3D: Shape Reconstruction From Single RGB Images for Applications in Transportation Logistics}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops}, month = {June}, year = {2023}, pages = {4402-4412} }</code></pre> <p> </p>
Deeptangle Dataset: Labelled Experimental and Synthetic Videos of Swimming and Overlapping C. elegans worms
<p>This repository contains the dataset employed in the paper <a href="https://arxiv.org/abs/2301.04460">Fast spline detection in high density microscopy data</a>.</p> <p>Three files are provided:</p> <p>1. <em>videos.zip</em>: raw experimental videos.<br> 2. <em>labeled_data.zip</em>: labelled sections of experimental videos used for evaluation<br> 3. <em>syntehthic_dataset.zip</em>: synthetic dataset used for training</p> <p><strong>Labelled data</strong></p> <p>Sections are named as VIDEONAME_FRAMENUMBER_SECTIONID.<br> All labels are in labels.json and correspond to the middle frame of the clip (05.png).<br> Example plotting script is provided (<em>plot_data.py</em>).</p> <p><strong>Synthetic data</strong></p> <p>Clips are named as NUMBEROFWORMS_ID.<br> Labels (for all frames) are stored in labels.npy.<br> Example plotting script is provided (plot_data.py).</p> <p> </p> <p><strong>Related</strong></p> <p>Paper: <a href="https://arxiv.org/abs/2301.04460">https://arxiv.org/abs/2301.04460</a></p> <p>Deeptangle code: <a href="https://github.com/kirkegaardlab/deeptangle">https://github.com/kirkegaardlab/deeptangle</a></p> <p>Labelling tool: <a href="https://github.com/kirkegaardlab/deeptanglelabel">https://github.com/kirkegaardlab/deeptanglelabel</a></p> <p>---</p> <p>If used, please cite</p> <p><em>Albert Alonso & Julius B. Kirkegaard. Fast spline detection in high density microscopy data. 2023.</em></p>
SPASS dataset: A synthetic polyphonic dataset with spatiotemporal labels of sound sources
<p>SPASS is a synthetic dataset that consists of 10-seconds audio segments from 5 acoustic scenes:</p> <ul> <li>Park</li> <li>Square</li> <li>Street</li> <li>Waterfront</li> <li>Market</li> </ul> <p>Each acoustic scene has 5,000 audio recordings and its corresponding metadata.</p> <p>The audio recordings were created using a 3D acoustic simulation environment (RAVEN, <a href="https://www.virtualacoustics.org/RAVEN/">https://www.virtualacoustics.org/RAVEN/</a>).</p> <p>SPASS was made as a training dataset for the FuSA system (<a href="https://www.acusticauach.cl/fusa/">https://www.acusticauach.cl/fusa/</a>). This is a polyphonic dataset for Sound Event Detection (SED) tasks.</p> <p>The metadata files includes the class of each sound event, their onset and offset in time, the position in the space (cartesian) and their final position if the class was moving.</p> <p>This research was funded by ANID FONDEF grant number ID20I10333.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.