Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
49
datasets available to search
ShareScore release 0.9.0
Dataset results
49 results for “Space Exploration”
Benchmark for deterministic traffic simulator - parameter space exploration (Prague, Jun 6 2021)
<p>The benchmark is meant for deterministic traffic simulator for optimising traffic flow within a city. The simulator is one part of a traffic modeling framework for intelligent transportation in smart cities. In contrast to standard navigation systems where the navigation is optimised for drivers, we aim to optimise a distribution of the global traffic flow. We utilise HPC resources for the simulator’s parameters exploration for which EVEREST SDK is used.</p> <p>The traffic simulator is available at: <a href="https://github.com/It4innovations/ruth">github.com/It4innovations/ruth</a><strong>.</strong></p> <p><br> The benchmark contains input data, routing map, and skript to run it with HyperQueue. Simulator v1.0 was used.</p>
Dataset linking to the paper "Exploring characteristics of national forest inventories for integration with global space-based forest biomass data"
<p>The dataset links to the study titled “Exploring characteristics of national forest inventories for integration with global space-based forest biomass data”. This study is published in the journal “Science of the Total Environment” and the publication can be found at <a href="https://doi.org/10.1016/j.scitotenv.2022.157788">https://doi.org/10.1016/j.scitotenv.2022.157788</a>. The dataset contains four csv files that were used to produce the results and other figures in the paper. The description of the individual data files contained in the dataset is given below.</p> <p><strong>NFI availability and characteristics data: </strong>The data file “NFI_availability_characteristics.csv” contains data on the total number of NFIs, the NFI extent, and the year of the most recent NFI in countries with NFI as reported in FRA 2020 country reports. The respective data variables in the data file are termed as Number_of_NFI, Latest_NFI_extent_FRA2020, and Latest_NFI_year_FRA2020 (NFI years generally refer to the years of data collection). In addition, the data file contains data on the region and tropical domain per country. The tropical and subtropical countries were considered tropical in the analysis and interpretation of the results. These data were used to produce Figure 2 of the study. ArcMap 10.7.1 was used for this purpose. </p> <p><strong>National biomass intercomparison data: </strong>The data file “national_biomass_intercomparison.csv” contains national forest AGB data for the year 2018 from FRA 2020 and CCI Biomass product that were used in the national biomass intercomparison analysis. The total (tons) and average space-based AGB (tons/ha) are extracted directly from the CCI Biomass Map 2018 for each country included in the study. The processing is done in Python and R environments. The spatial resolution of the map is 100 m. The average FRA AGB data in tons per ha was compiled from FRA 2020 country reports. The total FRA AGB data (tons) was estimated by multiplying each country's average FRA AGB data with FRA forest area data (in ha).</p> <p>The data unit for total AGB was converted from tons to gigaton (Gt) in intercomparison analysis. The total CCI Map AGB estimates used in the analysis are termed as CCI_MAP_AGB_Gt in the data file and the average as CCI_Map_AGB_tons.ha. Similarly, the total FRA AGB data are termed as FRA_AGB_Gt and the average as FRA_AGB_ton.ha. The NFI availability and temporality were also used in intercomparison analysis and this data is termed as Latest_NFI_year_FRA2020 in the data file. The data were used to produce Figure 3 of the study in the R environment.</p> <p><strong>NFI plot design characteristics: </strong>The data file named “NFI_plot_design_characteristics.csv” contains data on variables that were used in the analysis of NFI plot designs in 46 tropical countries. This data file mainly contains the data that was used to produce Figure 4 and Figure 6 in the R environment. The value “uniform” in the sampling_stratification variable means no stratification was used in the sampling design. The variable name “psu” stands for primary sampling unit (both cluster and single plots), “psu_distance_km” for the distance between primary sampling units in km, “cluster_plotdis_m” for the distance between plots in meter in the cluster, “plotsize_ha” for plot (single and cluster plots ) size in ha, “plotshape” for plot shapes (single and cluster plots), “ILUA” for Integrated Land Use Assessment. The data were compiled from the latest NFI design manuals and NFI reports.</p> <p><strong>NFI years: </strong>The data file “NFI_years_tropical_countries_data.csv” contains data on NFI years of the latest NFI in 46 tropical countries that were used to produce Figure 1 using ArcMap 10.7.1. The years generally refer to the last years of data collection. Data were compiled from the latest country NFI design manual or NFI report. This included both ongoing and completed NFI.</p>
Exploring chemical space in the search for improved Azoheteroarene-based photoswitches
<p>In the quest for improved photo switches, azoheteroarenes have emerged as a potential alternative to azobenzene. However, to date the number and types of these species that have subjected to study is insufficient to provide an in-depth understanding of the photochemical effects brought about by different substituents. Here, we computationally screen the optical properties and thermal stabilities of 512 azoheteroarenes that consist of eight different N-containing heteroarenes combined with 64 substitution patterns. The most promising compounds are identified and their properties rationalized based on the nature of the azoheteroarene core and the location and type of substitution patterns.</p>
Input geophysical and geological data for "Geologically constrained geometry inversion and null-space navigation to explore alternative geological scenarios: a case study in the Western Pyrenees"
<p>This is a companion dataset to the manuscript: <br><br>Geologically constrained geometry inversion and null-space navigation to explore alternative geological scenarios: a case study in the Western Pyrenees,</p><p>by: Jeremie Giraud , Mary Ford, Guillaume Caumon, Lachlan Grose, Vitaliy Ogarko, Roland Martin, and Paul Cupillard.<br><br>This dataset contains the input data used in the inversion, in terms of the gravity data and the geological data used in the inversion.<br><br>The *.txt file contains the gravity data as inverted in the manuscript: X, Y, Z, Value.<br>The *.csv file contains the geological data: location of the contacts and orientation data.</p>
Data and code from: SIDERITE: Unveiling Hidden Siderophore Diversity in the Chemical Space Through Digital Exploration
<h1>TMAP of COCONUT database</h1> <p>The script and data used in SIDERITE paper to generate TMAP picture (Figure S3 in supplementart material).</p> <p>Requirement: tmap</p> <p>You can install tmap by conda.</p> <blockquote> <p>conda create -n tmap python=3.7</p> <p>conda activate tmap</p> <p>conda install -c tmap tmap</p> <p>pip install faerun</p> <p>pip install matplotlib</p> <p>conda install scipy</p> <p>conda install -c rdkit rdkit</p> <p>conda install -c conda-forge mhfp</p> </blockquote> <p> </p> <p>Usage: python plot_COCONUT.py</p> <p>Then it will use COCONUT.csv to generate index.html and index.js. Open index.html to see result.</p> <h1>Other large input files</h1> <p>Tanimoto_COCONUT_SIDERITE.xlsx is the input file in <a href="https://github.com/RuolinHe/SIDERITE/blob/main/predicted_new/Clustering_coconut.m">SIDERITE/predicted_new/Clustering_coconut.m at main · RuolinHe/SIDERITE</a>.</p> <p> </p> <p>Sid_structure_output3.xlsx is used in <a href="https://github.com/RuolinHe/SIDERITE/blob/main/statistics/Figure1.m" target="_blank" rel="noopener">SIDERITE/statistics/Figure1.m at main · RuolinHe/SIDERITE</a>, <a href="https://github.com/RuolinHe/SIDERITE/blob/main/TAMP/tmap_code.m" target="_blank" rel="noopener">SIDERITE/TAMP/tmap_code.m at main · RuolinHe/SIDERITE</a> and <a href="https://github.com/RuolinHe/SIDERITE/blob/main/siderophore_process/Sid_process_code.m" target="_blank" rel="noopener">SIDERITE/siderophore_process/Sid_process_code.m at main · RuolinHe/SIDERITE</a>.</p> <p> </p> <p>COCONUT4MetFrag_Canonical.xlsx is the input file in <a href="https://github.com/RuolinHe/SIDERITE/blob/main/TAMP/tmap_code.m" target="_blank" rel="noopener">SIDERITE/TAMP/tmap_code.m at main · RuolinHe/SIDERITE</a>.</p> <p> </p> <p>COCONUT_r.txt is the output file in <a href="https://github.com/RuolinHe/SIDERITE/blob/main/TAMP/tmap_code.m" target="_blank" rel="noopener">SIDERITE/TAMP/tmap_code.m at main · RuolinHe/SIDERITE</a>. and the input file in <a href="https://github.com/RuolinHe/SIDERITE/blob/main/predicted_new/isSiderophore1.py" target="_blank" rel="noopener">SIDERITE/predicted_new/isSiderophore1.py at main · RuolinHe/SIDERITE.</a></p> <p> </p> <p>COCONUT4MetFrag.xlsx is the input file in <a href="https://github.com/RuolinHe/SIDERITE/blob/main/TAMP/CheckSMILES2.py" target="_blank" rel="noopener">SIDERITE/TAMP/CheckSMILES2.py at main · RuolinHe/SIDERITE</a>.</p> <p> </p>
Small dataset machine-learning approach for efficient design space exploration: engineering ZnTe-based high-entropy alloys for water splitting
<p>Atomic structure data used in the research article entitled "Small Dataset Machine-Learning Approaches to Explore the Design Space of High-Entropy Alloys: Engineering ZnTe-based Multicomponent Alloys for the Photo-Splitting of Water"</p>
IMU-Based Tip-Over Dataset for Space Exploration Rover Dynamics and Stability Analysis
<p>This dataset provides a comprehensive collection of Inertial Measurement Unit (IMU) sensor data captured from a space exploration rover under varying conditions of terrain, speed, and inclination. The primary goal of this dataset is to enable the study of dynamic stability, specifically the detection and analysis of tip-over events, which are critical for safe and efficient operation of autonomous rovers in extraterrestrial environments.</p> <p> </p> <p><em>This dataset is provided by the Robotics Innivation Center, DFKI GmbH.</em></p> <p><em>The grant was provided by Federal Ministry for Economic Affairs and Climate Action </em></p> <p><em>Grant number: 50RA2124</em></p>
CyZ: MARS Space Exploration Dataset.
<p>Images from NASA missions of the celestial body.</p> <p>Repository: https://github.com/decurtoidiaz/cyz</p> <p>Authors:</p> <p>J. de Curtò c@decurto.be</p> <p>I. de Zarzà z@dezarza.be</p> <p>------------------------------------------<br> File Information from CyZ-1.1<br> ------------------------------------------</p> <ul> <li>Curiosity (cyz/curiosity_cyz). <ul> <li>png (cyz/curiosity_cyz/png). <ul> <li>PNG files for the corresponding cameras.</li> </ul> </li> <li>csv (cyz/curiosity_cyz/csv). <ul> <li>CSV files. </li> </ul> </li> </ul> </li> <li>Perseverance (cyz/perseverance_cyz). <ul> <li>png (cyz/perseverance_cyz/png). <ul> <li>PNG files for the corresponding cameras. </li> </ul> </li> <li>csv (cyz/perseverance_cyz/csv). <ul> <li>CSV files. </li> </ul> </li> </ul> </li> </ul>
Research data for "Exploring the configurational space of amorphous graphene with machine-learned atomic energies"
<p>This dataset supports the paper: "Exploring the configurational space of amorphous graphene with machine-learned atomic energies" (<a href="https://doi.org/10.1039/D2SC04326B">https://doi.org/10.1039/D2SC04326B</a>).</p> <p>Trajectory data for the 200-atom structures (Fig. 3) and the final configurations for the 612-atom structures as well as the GAP-17-optimised 610-atom structure from Toh et al are provided (Fig. 4). Additionally, the structures used for data analysis in Fig. 5 are given.</p> <p>The files are in extended xyz (.xyz) format and contain the raw data for coordinates, forces, and atomic energies (labelled 'c_1'). The files also contain the atomic energies relative to pristine graphene, labelled "Energy_per_atom", and the locally averaged energy relative to pristine graphene, labelled "NN_Energy_per_atom". Topological information is included at the end of the .xyz file for the 612-atom structures ('fig_4'/) and for the structures in 'fig_5/'.</p> <p>All raw atomic energies were computed using LAMMPS default settings and were output with six significant figures, with the exception of the Toh et al. structure (for which ASE was used, outputting a higher number of significant figures). </p> <p>The data can be read using, for example, the Atomic Simulation Environment (ASE), or visualised using Ovito.</p> <p> </p>
A DSL based toolchain for design space exploration in structured parallel programming
<p>We introduce a DSL based toolchain supporting the design of parallel applications where parallelism is structured after parallel design pattern compositions. A DSL provides the possibility to write high level parallel design pattern expressions representing the structure of parallel applications, to refactor the pattern expressions, to evaluate their non-functional properties (e.g. ideal performance, total parallelism degree, etc.) and finally to generate parallel code ready to be compiled and run on different target architectures. We discuss a proof-of-concept prototype implementation of the proposed toolchain generating FastFlow code and show some preliminary results achieved using the prototype implementation.</p>
Sequence data for 'Machine-driven parameter-space exploration of biochemical reactions'
<p>The development of complex, multi-step <em>omics</em> methods in molecular biology is a laborious, costly, iterative and often intuition-bound process where an optimum is sought in a parameter space through step-by-step optimisations. The the difficulty of miniaturising assays and the cost of the experiments limit the dynamic range and the number of parameters that can be explored. However, because of non-linearities of the response of biochemical systems to their reagent concentrations, a broad dynamic range is necessary. Here we demonstrate the use of a high-performance nanoliter handling platform (Labcyte Echo 525) and computer generation of liquid transfer programs to explore in quadruplicates more than 600 combination of 4 parameters of a biochemical reaction, which lead us to uncover non-linear responses, parameter interactions and novel mechanical insights. With the increased availability of « <em>cloud biology</em>» computer-driven laboratory platforms, our results participate in changing methods development for biotechnology towards reproducible, computer-aided exhaustive characterisation of biochemical systems.</p> <p>This dataset contains the raw sequencing data produced with an Illumina MiSeq instrument for this project. FASTQ files and sample sheets are found in the usual location (Data/Intensities/BaseCalls). The "Thumbnail_Images" and "L001" directories were deleted to save space.</p> <p>Run IDs: 171227_M00528_0321_000000000-B4GLP, 180403_M00528_0348_000000000-B4GP8, 180517_M00528_0364_000000000-BRGK6, 180123_M00528_0325_000000000-B4PCK, 180411_M00528_0351_000000000-BN3BL, 180606_M00528_0367_000000000-BN3FG, 180326_M00528_0346_000000000-B4GJR, 180501_M00528_0359_000000000-B4PJY, 180607_M00528_0368_000000000-BN9KM</p> <p> </p>
Data from: Exploring thermal tolerance across time and space in a tropical bivalve, Pinctada margaritifera
Open the record for dataset details and reuse information.
Reward-based option competition in human dorsal stream and transition from stochastic exploration to exploitation in continuous space
<p>Primates exploring and exploiting a continuous sensorimotor space rely on dynamic maps in the dorsal stream. Two complementary perspectives exist on how these maps encode rewards. Reinforcement learning models integrate rewards incrementally over time, efficiently resolving the exploration/exploitation dilemma. Working memory buffer models explain rapid plasticity of parietal maps but lack a plausible exploration/exploitation policy. The reinforcement learning model presented here unifies both accounts, enabling rapid, information-compressing map updates and efficient transition from exploration to exploitation. As predicted by our model, activity in human fronto-parietal dorsal stream regions, but not in <em>MT+</em>, tracks the number of competing options, as preferred options are selectively maintained on the map while spatiotemporally distant alternatives are compressed out. When valuable new options are uncovered, posterior beta<sub>1</sub>/alpha oscillations desynchronize within 0.4-0.7 s, consistent with option encoding by competing beta<sub>1</sub>-stabilized subpopulations. Altogether, outcomes matching locally cached reward representations rapidly update parietal maps, biasing choices toward often-sampled, rewarded options.</p>
Efficiency and suitability when exploring the conformational space of phase transfer catalysts - Supporting information
<p>Supporting information for Efficiency and suitability when exploring the conformational space of phase transfer catalysts. In this study, a complete exploration of the conformational space within phase transfer catalysts by means of computational methods benchmarking is presented. For this particular research work, only the most applied conformational analysis approaches have been chosen to characterise different Cinchona alkaloid-based phase transfer catalysts. This particular benchmarking study aims to rigorously compare the performance of different conformational methods, determining the strengths of each method and to provide recommendations regarding suitable choices of methods for an analysis. <strong> </strong></p>
Supplementary Data for "Exploring the Chemical Space of Metal–Organic Frameworks with rht Topology for High Capacity Hydrogen Storage"
<p>This dataset includes optimization and simulation inputs, optimized structures, building blocks, computed hydrogen uptakes and textural properties of the metal–organic frameworks that correspond with work in "Exploring the Chemical Space of Metal–Organic Frameworks with rht Topology for High Capacity Hydrogen Storage" (DOI: 10.1021/acs.jpcc.4c00638).</p>
Supporting files for "Simulation-Guided Conformational Space Exploration to Assess Reactive Conformations of a Ribozyme."
<p>The archived folder contains the trajectory reported in the paper "Simulation-Guided Conformational Space Exploration to Assess Reactive Conformations of a Ribozyme" by S. Forget, E. Duboue-Dijon, G. Stirnemann, <em>J. Chem. Theory Comput.</em> 2024, as well as all the input and force field files necessary to repeat the simulations, using the Gromacs software.</p> <p>The Subdirectory "Procedures" contains:</p> <ul> <li>the explicit descriptions of the equilibration procedures with the mdp parameter files of each equilibration steps.</li> <li>3 scripts which are run to generate the replicas of REST2 simulations.</li> </ul> <p>For any additional information, the authors can be contacted by email: <a rel="noreferrer">selene.forget@gmail.com</a> and guillaume.stirnemann@ens.psl.eu</p>
PubChem and ChEMBL-series processed dataset used in Exhaustive local chemical space exploration using a transformer model
<p>PubChem and ChEMBL-series processed dataset used in <span>Exhaustive local chemical space exploration using </span><span>a transformer model</span></p>
Data for: Exploring Transition States of Protein Conformational Changes via Out-of-Distribution Detection in the Hyperspherical Latent Space
<p>This contains all the TS-DART training results as well as the raw MD simulation data reported in the preprint "Exploring Transition States of Protein Conformational Changes via Out-of-Distribution Detection in the Hyperspherical Latent Space".</p>
DAPI images, molecules and segmentation boundaries for: A Spatiotemporal Atlas of Mouse Gastrulation and Early Organogenesis to Explore Axial Patterning and Project In Vitro Models onto In Vivo Space
<div> </div> <p><strong>Data Description</strong></p> <ol> <li><strong>Stitched & rotated DAPI images</strong> - tiff file format filename indicates sample and optical z-slice position, i.e. embryo3_z5.tif is the DAPI image for embryo 3 in optical z-slice 5. Also provided in PNG format.</li> <li><strong>Detected molecules and cell segmentation in MoleculeExperiment objects</strong> - RDS files to read data using the MoleculeExperiment format in R/Bioconductor. Filename embryo3_z5.Rds indicates MoleculeExperiment RDS file for embryo 3 in optical z-slice 5. Coordinates are provided in microns. Note that z-slices 2 and 5 are only provided for embryos 1,2,3 as they were originally provided in Lohoff et al, Nature Biotechnology, 2023.</li> <li><strong>Pixels-to-microns conversion</strong> - pixelSize.R Simple R script/text to indicate the size of each pixel in the DAPI images, this is to align the coordinate systems between the DAPI images and molecules.<br><br> <div> <h4>Project Abstract</h4> </div> <p>At the onset of murine gastrulation, pluripotent epiblast cells migrate through the primitive streak, generating mesodermal and endodermal precursors, while the ectoderm arises from the remaining epiblast. Together, these germ layers establish the body plan, defining major body axes and initiating organogenesis. Although comprehensive single cell transcriptional atlases of dissociated mouse embryos across embryonic stages have provided valuable insights during gastrulation, the spatial context for cell differentiation and tissue patterning remain underexplored. In this study, we employed spatial transcriptomics to measure gene expression in mouse embryos at E6.5 and E7.5 and integrated these datasets with previously published E8.5 spatial transcriptomics and a scRNA-seq atlas spanning E6.5 to E9.5. This approach resulted in a comprehensive spatiotemporal atlas, comprising over 150,000 cells with 88 refined cell type annotations as well as genome-wide transcriptional imputation during mouse gastrulation and early organogenesis. The atlas facilitates exploration of gene expression dynamics along anterior-posterior and dorsal-ventral axes at cell type, tissue, and organismal scales, revealing insights into mesodermal fate decisions within the primitive streak. Moreover, we developed a bioinformatics pipeline to project additional scRNA-seq datasets into a spatiotemporal framework and demonstrate its utility by analysing cardiovascular models of gastrulation3. To maximise impact, the atlas is publicly accessible via a user-friendly web portal empowering the wider developmental and stem cell biology communities to explore mechanisms of early mouse development in a spatiotemporal context.</p> </li> </ol>
Genome alignments for 'Machine-driven parameter-space exploration of biochemical reactions'
<p>The development of complex, multi-step <em>omics</em> methods in molecular biology is a laborious, costly, iterative and often intuition-bound process where an optimum is sought in a parameter space through step-by-step optimisations. The the difficulty of miniaturising assays and the cost of the experiments limit the dynamic range and the number of parameters that can be explored. However, because of non-linearities of the response of biochemical systems to their reagent concentrations, a broad dynamic range is necessary. Here we demonstrate the use of a high-performance nanoliter handling platform (Labcyte Echo 525) and computer generation of liquid transfer programs to explore in quadruplicates more than 600 combination of 4 parameters of a biochemical reaction, which lead us to uncover non-linear responses, parameter interactions and novel mechanical insights. With the increased availability of « <em>cloud biology</em> » computer-driven laboratory platforms, our results participate in changing methods development for biotechnology towards reproducible, computer-aided exhaustive characterisation of biochemical systems.</p> <p>This dataset contains the sequence alignments and other processing files produced by running the raw data (10.5281/zenodo.1680999) through a processing pipeline using the MOIRAI workflow manager. The most important output is the "CAGEscan_fragments" directories and represent the alignment of single mRNA molecules, which can be further analysed using the "CAGEr" software package available from Bioconductor.</p> <p>Run IDs: 171227_M00528_0321_000000000-B4GLP, 180123_M00528_0325_000000000-B4PCK, 180326_M00528_0346_000000000-B4GJR, 180403_M00528_0348_000000000-B4GP8, 180411_M00528_0351_000000000-BN3BL, 180501_M00528_0359_000000000-B4PJY, 180517_M00528_0364_000000000-BRGK6 180606_M00528_0367_000000000-BN3FG, 180607_M00528_0368_000000000-BN9KM</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.