Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
453
datasets available to search
ShareScore release 0.7.1
Dataset results
453 results for “process model”
Nearshore high-frequency temporal water quality observations and process-based modeling of aquatic ecosystem metabolism in Lake Tahoe completed by members of the Blaszczak Lab at the University of Nevada Reno, 2021-2023
The overarching goal of this project was to develop a process-based understanding of how watershed-to-lake connections drive nearshore productivity dynamics in a large oligotrophic mountain lake (Lake Tahoe). We addressed this goal through a combined approach of high-frequency sensor deployment and maintenance, ecosystem metabolism modeling, laboratory incubations, and routine monitoring of water chemistry and other parameters. The data we collected as part of this project and the ecosystem metabolism estimates we generated demonstrate how variable ecosystem productivity is in time and space in the nearshore of Lake Tahoe. Although maintenance of the sensor arrays during the exceptional winter of 2023 was challenging, we were able to capture the data necessary to estimate a complete time series of metabolic activity across two years with very different hydroclimatic conditions. Throughout this project we accomplished the following: 1. We generated over two years of daily estimates of ecosystem metabolism (gross primary productivity, ecosystem respiration, and net ecosystem productivity) from multiple locations on both the east and west shores of the lake and from areas in close proximity to and far away from stream water inflows. 2. We measured ammonium (NH4+) and nitrate (NO3-) concentrations in surface water samples from both Glenbrook and Blackwood creeks and the nearshore of Lake Tahoe for over two years. 3. We quantified rates of NH4+ and NO3- uptake in benthic samples of the dominant substrate type collected during peak streamflow, the receding limb, and baseflow conditions in 2023 from multiple locations in the nearshore using established laboratory incubation methods. 4. Finally, we used a combination of time series models and structural equation modeling to integrate our results and improve understanding of the direct and indirect effects of hydroclimatic variability on observed patterns in ecosystem metabolism in the nearshore. See this git code repository
Larval transport pathways from three prominent sand lance habitats in the Gulf of Maine: otolith data, model data, and post-processed model data products
This dataset includes hatch and larval period for sand lance collected in 2019 and results from particle tracking runs of simulated sand lance larvae throughout the Northeast U.S. Shelf as part of Long-Term Ecological Research (NES-LTER). Release dates vary by region, corresponding to hatch and settlement dates of settling sand lance collected in 2019. Particles were depth-keeping throughout the upper 40 m to best replicate our understanding of the vertical distribution of sand lance larvae. Data were used to determine the average particle transport pathways from these sand lance habitats, including connectivity among the three hotspots, and spatial variability of connectivity within each hotspot. Further information can be found within the manuscript: Suca, J. J., Ji, R., Baumann, H., Pham, K., Silva, T. L., Wiley, D. N., Feng, Z., & Llopiz, J. K. (2022). Larval transport pathways from three prominent sand lance habitats in the Gulf of Maine. Fisheries Oceanography, 31( 3), 333-352. https://doi.org/10.1111/fog.12580
Numerical weather simulation using COSMOiso in June 2019 during L-WAIVE field campaign: selected model output and post-processed data.
<p>This dataset consists of extracts from a simulation with the isotope-enabled regional numerical weather prediction model COSMOiso, which covers the timespan of the Lacustrine-Water vApor Isotope inVentory Experiment (L-WAIVE) field campaign taking place in June 2019 in the Annecy valley in the French Alps (Chazette et al. 2021).The simulation has a horizontal resolution of 0.1° (~10km) and 40 vertical levels.</p><p>This COSMOiso simulation is used in Thurnherr et al. (submitted) to compare stable water isotope measurements from various platforms. Here, we provide selected model outputs and post-processed data used in this comparison study. The post-processed data contain:</p><ol><li>COSMOiso output files for time steps 20190612_12, 20190613_12, 20190615_13, 20190616_13, 20190617_12, 20190622_12.</li><li>Pressure weighted total and subcolumn averages for time steps 20190612_12, 20190613_12, 20190615_13, 20190616_13, 20190617_12, 20190622_12.</li><li>Vertical cross section of selected variables at Annecy, the location of the L-WAIVE field campaign, for the simulation time window.</li><li>Interpolated time series of subcolumn and total column averages at Annecy, the location of the L-WAIVE field campaign, for the simulation time window.</li><li>Interpolated variables along the flight tracks from the L-WAIVE campaign (see Sodemann and Seidl, 2023).</li></ol><p>See also README files for more details on the provided data.</p><p>To access further model output and post-processed data, please contact the dataset authors.</p>
Dataset: An Analytic Hierarchy Process-Based Multicriteria Model for Component Selection in a Computational Numerical Control (CNC) Machine
<p><i><strong>"An Analytic Hierarchy Process-Based Multicriteria Model for Component Selection in a Computational Numerical Control (CNC) Machine"</strong></i></p><p><i>CHILECON 2023 - </i><a href="https://site.ieee.org/chilesur/ieee-chilecon-2023/"><i>https://site.ieee.org/chilesur/ieee-chilecon-2023/</i></a><i> </i></p><p>---</p><p>En el marco del trabajo de referencia, los autores ponemos a disposición de los lectores la base de datos utilizada para el proceso de toma de decisión multicriterio para la selección del software y del MCU de una maquina CNC. </p><p>En el repositorio podrán encontrar los datos referentes a los criterios, subcriterios, indicadores, datos, fuentes de los datos extraídos, política de decisión, cálculos de las evaluaciones de los modelos AHP aplicados y el análisis de sensibilidad de estos. Además, podrán encontrar las gráficas utilizadas en el estudio en la mejor calidad posible. </p><p>El material fue puesto a disposición de todos los interesados para fines académicos y científicos. </p><p>Atte. </p><p>Los autores. </p><p>---</p>
Processing of 3-D Polygon Mesh Model and Radio Propagation Simulations in a Cave: Surface Reconstruction from Point Cloud, Simplification of the Mesh, and Ray Tracing
<p><strong>ABOUT</strong></p><p>This repository includes mesh data from cave geometry scanning and processing, and radio propagation data from ray tracing simulations.</p><p>The geometry data is obtained with laser scanning in a cave in Slovenija. </p><p>The geometry processing includes (i) 3-D shape reconstruction - surface reconstruction from point cloud data and (ii) simplification - reduction of the geometric complexity of the 3-D mesh model. </p><p>The radio propagation data is obtained using CloudRT [1] ray-tracing simulator. </p><p>The obtained propagation-related quantities include information about the propagation mechanism, interactions with the geometry, received power, delay, azimuth and elevation angles of arrival and departure, and path loss. </p><p> </p><p><strong>AUTHORS</strong></p><p>Teodora Kocevska, Andrej Hrovat, Tomaž Javornik</p><p>Department of Communication Systems</p><p>Jožef Stefan Institute, SI-1000 Ljubljana, Slovenia</p><p>teodora.kocevska@ijs.si</p><p> </p><p><strong>GEOMETRY PROCESSING</strong></p><p>The cave segment used for the propagation calculations is selected from a point cloud obtained in a cave in Litia, Slovenia. The point cloud is obtained with 3-D laser scanning of the environment. The selected segment is approx. 58 m long. Several parameter configurations were considered for 3-D shape reconstruction, including Poisson surface reconstruction with octree depths of 8, 10, and 12. Geometries that represent the cave shape and have different levels of complexity were created and studied. In the simplification process, one and two-stage simplification was explored using the Quadric Edge Collapse Decimation approach. </p><p> </p><p><strong>RADIO SETUP</strong></p><p>The transmitter (Tx) is fixed at the entrance of the cave and the receiver (Rx) is moved along the cave in 40 positions with a step of 1 m.</p><p>Omnidirectional antennas at the Tx and Rx sites and vertical polarization are considered. The antenna is mounted 1.5 m above the ground.</p><p>The start frequency is 3.5 GHz, the end frequency is 3.6 GHz and the step is 10 MHz. Direct propagation and first-order reflection are considered. </p><p>The cave geometry is represented by a triangular mesh, and the material of the cave is wet earth. The material electromagnetic properties are selected according to the specifications presented in [2].</p><p> </p><p><strong>FOLDER STRUCTURE</strong></p><p>The folder structure is:</p><p> - Polygon_Mesh_Models</p><p> <i># 3-D environment models with varying </i>levels<i> of geometry complexity</i></p><p> - Reconstruction_Segmen1_Poisson_Surface_Reconstruction</p><p> - Simplification_Segment1_Quadric_Edge_Collapse_Decimation</p><p> - Propagation_Data</p><p> <i># Propagation quantities of all rays between a transmitter and receiver</i></p><p> - AllRay_PropData</p><p> - PathLoss</p><p> - readme.txt</p><p> - RayTracing_EnvironmentModel</p><p> <i> # Final environment model used for ray tracing simulations</i></p><p> - Cave_MeshModel.json</p><p> - Cave_MeshModel.skb</p><p> - Cave_MeshModel.skp</p><p> - RayTracing_MaterialProperties</p><p> <i># Properties of the materials in the environment</i></p><p> - materials.json</p><p> - materials.mtl</p><p> - readme.txt</p><p> - Cave_Length.txt</p><p> <i># Length between selected locations in the environment</i></p><p> - Cave_Segment1_visual.png</p><p> <i> # Visualization of the environment segment used for propagation calculation</i></p><p> - readme.txt</p><p> <i># Overall description </i></p><p><strong>REFERENCES</strong></p><p>[1] D. He, B. Ai, K. Guan, L. Wang, Z. Zhong, and T. Kürner, "The Design and Applications of High-Performance Ray-Tracing Simulation Platform for 5G and Beyond Wireless Communications: A Tutorial," in IEEE Communications Surveys & Tutorials, vol. 21, no. 1, pp. 10-27, First quarter 2019, doi: 10.1109/COMST.2018.2865724.</p><p>[2] R. sector of International Telecommunication Union (ITU-R), "Effects of building materials and structures on radio wave propagation above about 100 MHz," International Telecommunication Union, ITU-R Recommendation P.2040-2, 2021.</p><p> </p><p><strong>ACKNOWLEDGEMENT</strong></p><p>This work was supported by the Slovenian Research Agency under grant <strong>J2-3048</strong>.</p><p> </p>
Evaluation datasets and results for the paper "Enhancing Business Process Simulation Models with Extraneous Activity Delays"
<p>Event-logs and Business Process Simulation Models used in the experimentation of the paper "Enhancing Business Process Simulation Models with Extraneous Activity Delays", where the '<em>inputs</em>' folder contains all the files used as input, and the '<em>output</em>' folder the results of the evaluation.</p> <p> </p> <p><em><strong>Inputs</strong></em>: event-logs, BPS models, and simulation parameters used as input in the experimentation.</p> <ul> <li><em><strong>Real-life</strong></em>: real-life event logs, corresponding to two disjoint subsets of traces from an Academic Credentials' process, and the BPIC 2012 and BPIC 2017 event logs (filtered as explained in the paper), and the BPS model (plus simulation parameters) used as input for each dataset in the presented approach.</li> <li><em><strong>Synthetic</strong></em>: simulated event-logs and corresponding BPS models (plus simulation parameters) for four different processes with 0, 1, 3 and 5 timer events.</li> </ul> <p><em><strong>Outputs</strong></em>: results of the experimentation.</p> <ul> <li><em><strong>Real-life</strong></em>: results corresponding to the evaluation with real-life event logs. Each of the folders is composed by the original and the enhanced BPS models, 10 event logs simulated with each of them, two folders with the best iteration of the two hyperparameter optimization processes, and the values for the injected timers in each case. In addition, a CSV file with the EMD metrics (cycle time and absolute hour event distribution) for each dataset is provided.</li> <li><em><strong>Synthetic</strong></em>: results corresponding to the simulated event-logs. <ul> <li>Before-After: BPS models and discovered timer events for the four synthetic processes, with five timers placed before and after different activity instances.</li> <li>Complete: BPS models and quality measures (precision, recall, and SMAPE of the discovered timers) for the four synthetic processes with zero, one, three, and five timer events.</li> <li>Individual: event logs enhanced with the discovered extraneous delay for each activity instance, for the four synthetic processes with zero, one, three, and five timer events; and SMAPE of the estimations.</li> </ul> </li> </ul>
Modified WRF/Chem source code, output data, and post-processing scripts for the GMD manuscript "Evaluation of WRF/Chem model (v3.9.1.1) real-time air quality forecasts over the Eastern Mediterranean"
<p>Here you will find the modified WRF/Chem code used in the simulations, the scripts used for post-processing and the model output data used in the manuscript. </p> <p>Two modifications have been made in module_aerosols_soa_vbs.F:</p> <ol> <li>ch_dust is set to1.0D-9*0.36</li> <li>The model is set not to initialize during restarts</li> </ol> <p>The model data directory includes:</p> <ol> <li>Two csv files (Winter and Summer) with the hourly concentrations of atmospheric pollutants at the locations of the ground stations. These data were used to produce Figures 4-8 in the manuscript as well as all the metrics.</li> <li>Two netcdf files (Winter and Summer) with the average ground concentrations of atmospheric pollutants over Cyprus. These data were use to produce Figure 3 in the manuscript. </li> </ol>
Process modeling, environmental and economic sustainability of the valorization of whey and eucalyptus residues for resveratrol biosynthesis
<p>Tables included in the article "Process modeling, environmental and economic sustainability of the valorization of whey and eucalyptus residues for resveratrol biosynthesis"</p>
Gaussian Process Model and Sensor Placement for Detroit Green Infrastructure: Datasets and Code
<ol> <li><strong>code.zip: </strong>Zip folder containing a folder titled "code" which holds: <ol> <li>csv file titled "MonitoredRainGardens.csv" containing the 14 monitored green infrastructure (GI) sites with their design and physiographic features;</li> <li>csv file titled "storm_constants.csv" which contain the computed decay constants for every storm in every GI during the measurement period;</li> <li>csv file titled "newGIsites_AllData.csv" which contain the other 130 GI sites in Detroit and their design and physiographic features;</li> <li>csv file titled "Detroit_Data_MeanDesignFeatures.csv" which contain the design and physiographic features for all of Detroit;</li> <li>Jupyter notebook titled "GI_GP_SensorPlacement.ipynb" which provides the code for training the GP models and displaying the sensor placement results;</li> <li>a folder titled "MATLAB" which contains the following: <ol> <li>folder titled "SFO" which contains the SFO toolbox for the sensor placement work</li> <li>file titled "sensor_placement.mlx" that contains the code for the sensor placement work</li> <li>several .mat files created in Python for importing into Matlab for the sensor placement work: "constants_sigma.mat", "constants_coords.mat", "GInew_sigma.mat", "GInew_coords.mat", and "R1_sensor.mat" through "R6_sensor.mat"</li> <li>several .mat files created in Matalb for importing into Python for visualizing the results: "MI_DETselectedGI.mat" and "DETselectedGI.mat"</li> </ol> </li> </ol> </li> </ol>
A CO2 valorization plant to produce light hydrocarbons: kinetic model, process design and life cycle assessment
<p>Supplementary material: Reaction indexes, Conservation equations, boundary conditions and used coefficients. Additional experimental results, Experimental data fitting, Stream properties and composition of the CO2 plant, Life Cycle Assessment indicators, assumptions and data input </p>
Modelling impacts of tramlines on soil erosion processes at the catchment scale.
<p>The data refers to the following article:</p> <p>Saggau, P., M. Kuhwald, W. B. Hamer, R. Duttmann (2021): Are compacted tramlines underestimated features in soil erosion modelling? A catchment‐scale analysis using a process‐based soil erosion model. In: Land Degradtaion & Development. doi: 10.1002/ldr.4161</p> <p>The data only contains distributable raw data and R-codes.</p>
Data, scripts, and figures of the article: Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models
<p>This data set contains the data, JMP scripts, and figures of the article titled "Processing weights of chickens determined by Dual-Energy X-Ray Absorptiometry. 2. Developing prediction models" to be published in the journal Animal - Open Space.</p>
Processed model output of the climate simulation in the study: The effects of diachronous surface uplift of the European Alps on regional climate and the isotopic composition of precipitation (δ18Op) [Boateng et al.]
<p><strong>The geodynamic evolution of the Alps suggests that the Alps did not rise monotonically due to the different post-collisional processes such as slab break-off. However, understanding such subsurface dynamics would require adequate knowledge about its surface uplift history. Stable isotope paleoaltimetry methods are widely used to infer past surface elevation using geologic archives. However, its accurate interpretation relies on attributing the extracted isotopic signal from proxies to surface uplift despite other influences such as climate. To resolve this issue, topographic sensitivity experiments across the Alps are used to investigate the impacts of the diachronous surface uplift on regional climate and δ18Op. The Atmospheric General Circulation Model ECHAM5 with water isotope tracking capabilities (ECHAM5-wiso) is used to simulate the climate with varied topographic scenarios. We present the processed (long-term means) model output of the relevant climate variables (i.e δ18Op, near-surface temperature, precipitation amount, near-surface meridional and zonal winds, mean sea level pressure, and elevation) in response to the changes in topography. The file names are representative of the topographic scenarios used for the simulations. For example, the file “W2E1.nc” is the model output produced by a topographic scenario in which the topography across the west-central Alps was set to 200% of its modern height, and the Eastern Alps were kept at 100%. The “CTL.nc” file contains model output from the control simulation that uses present-day topography. The datasets for instance can be used to select far-field sampling points for the δ-δ paleoaltimetry method that are not significantly affected by the topographic changes.</strong></p>
Multiscale continuum figures from Tratnyek et al. (2017) "In silico environmental chemical science: Properties and processes from statistical and computational modelling"
<p>Accessible versions of selected figures from Tratnyek et al. (2017) "In silico environmental chemical science: Properties and processes from statistical and computational modelling" Environ. Sci. Processes Impacts 19(3): 188-202. DOI: 10.1039/C7EM00053G.</p> <p>The Abstract Art figure shows a classification of variables for predictive/diagnostic models used in silico environmental chemical science, in terms of system scales and variable types. Figure 3 shows a continuum of system scales encompassing the whole scope of predictive/diagnostic modelling for in silico environmental chemical sciences, juxtaposing earth and biological scales.</p> <p>The published version of Figure 3 is tall, for two-column page-layouts, but a wide version of Figure 3 is provided for landscape oriented formats. The 300 dpi versions of each figure should be adequate resolution for most purposes, and therefore are recommended. The large versions of the figures may take significant time to download, but may be useful for high resolution applications.</p> <p>This work is from the perspectives/review paper at the beginning of a themed issue on "Quantitative Structure-Activity Relationships (QSARs) and Computational Chemistry Methods in the Environmental Chemical Sciences", published in the March 2017 issue of the Royal Society of Chemistry journal Environmental Sciences: Process and Impacts. The whole collection of papers can be accessed at rsc.li/qsars.</p>
Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline
<p>This upload contains the HZV029 Plasma and HZV029 Two-Phase dataset for reviewers of the "Data for common data models to streamline metabolomics processing and annotation, and implementation in a Python pipeline" submission. </p> <p>Both datasets will be uploaded to metabolomics workbench and the upload completed before final publication of the manuscript. For the he HZV029 Plasma datasets only the final run is included for any sample (i.e., failed injections or other samples with data quality issues that were reran during acquisition were omitted).</p> <p>Also included in the upload is the source code for the MetDataModel and the pcpfm at the time of manuscript re-submission and the pcpfm itself. If you find this upload in the future, please check out the github repos for more updated versions:</p> <p>https://github.com/shuzhao-li-lab/PythonCentricPipelineForMetabolomics</p> <p>https://github.com/shuzhao-li-lab/metDataModel</p> <p>The github repo does not store the input the data for space reasons, they only have the notebooks. However, the .zip here has both the notebooks by themselves in the notebook subdirectory and a separate directory with the notebooks and the data used to generate all the figures and results in the manuscript.</p> <p><strong>Some information that is needed to rerun this analysis:</strong></p> <p>Sequence files are critical to the functioning of the pipeline. The sequence files for all analyses are provided under sequence_files.zip. These can be used to recapitulate the analysis by eitehr changing the filepath to each acquisition to where you put it on your sytem or by placing the sequence file in the same directory as the mzml or raw. In the latter case, the pipeline will search for filenames matching the sample names. The sequence files also store some sample metadata such as the type of sample a given acquisition is (unknown, pooled, qc, etc...)</p> <p>.raw to .mzML conversion works well on MacOS but may not work well on other systems. You will need to use the ability to specify your own conversion command or convert files outside of the pipeline. </p> <p>To replicate the results, you do need to have the annotation sources downloaded which can be done using the pipeline. MS2 annotation requires the files in the AcquireX directory which is MS2 acquisitions on pooled HZV029 plasma samples.</p> <p>For the comparison between MetaboAnalystR and the pcpfm, subsets of the datasets were used. These subsets and the sequence files are in Subsets_for_performance_testing.zip. The sequences are also in the sequence_files directory as well</p> <p>The notebooks reference data in the analysis folders. Copies of these files are located with the notebooks to ease reproduction of the exact results in the paper; however, to do so, you will need to change paths to this data in the notebook. This lets the notebooks be ran during a rerun without copying intermediates back and forth and it keeps the github repo clean.</p> <p><strong>Version History:</strong></p> <p>This version is after reviewer comments and is for resubmission.</p> <p> </p> <p><strong>Contributions:</strong></p> <p>Joshua M Mitchell implemented the pipeline and was first author on the manuscript. Shuzhao Li is the corresponding author on the manuscript. </p> <p>Maheshwor Thapa performed the experiments to collect the HZV029 data. Yuanye Chi helped with testing and documenting the pipeline. </p> <p>Jiangou (Jeff) Xia and Zhiqiang Pang provided the R portion of the analysis. </p>
Supplementary Data from, "Causal health impacts of power plant emission controls under modeled and uncertain physical process interference."
<p>These data are used to conduct the analysis in, "<a href="https://arxiv.org/abs/2306.05665">Causal health impacts of power plant emission controls under modeled and uncertain physical process interference</a>," by Wikle and Zigler (2024), to appear in <em>Annals of Applied</em> Statistics. This is purely for archival purposes to facilitate access to and replication of the aforementioned analysis. Data were obtained from the following sources:</p> <ol> <li> U.S. Emissions Data [<a href="https://ampd.epa.gov/ampd">U.S. EPA, Air markets program data (AMPD)</a>] <ul> <li>AMPD_Unit_with_Sulfur_Content_and_Regulations_with_Facility_Attributes.csv</li> </ul> </li> <li> US Census 2016 American Community Survey [<a href="https://www.census.gov/programs-surveys/acs">US Census Bureau ACS</a>] <ul> <li>Census_2016_TxZCTA.RDS</li> <li><em>Note: data were obtained using the r package ‘<a href="https://walker-data.com/tidycensus/">tidycensus</a>’.</em></li> </ul> </li> <li> Daymet Annual Climate Summaries [<a href="https://daac.ornl.gov/DAYMET/guides/Daymet_V4_Annual_Climatology.html">Daymet Version 4</a>] <ul> <li>daymet_v4_prcp_annttl_na_2016.nc</li> <li>daymet_v4_tmax_annavg_na_2016.nc</li> <li>daymet_v4_tmin_annavg_na_2016.nc</li> <li>daymet_v4_vp_annavg_na_2016.nc</li> </ul> </li> <li> SO<sub>4</sub> and Black Carbon Concentrations [<a href="https://sites.wustl.edu/acag/datasets/surface-pm2-5/#V4.NA.03">Randall Martin Atmospheric Composition Analysis Group, North American Regional Estimates, version V4.NA.02</a>] <ul> <li>GWRwSPEC_BC_NA_201601_201612.nc</li> <li>GWRwSPEC_SO4_NA_201601_201612.nc</li> </ul> </li> <li> HyADS Coal-Attributed PM2.5 Concentrations [<a href="https://doi.org/10.1097/EDE.0000000000001024">Henneman et al. (2019)</a>] <ul> <li>HyADS_grids_pm25_byunit_2016.fst</li> <li>HyADS_grids_pm25_total_2016.fst</li> </ul> </li> <li> Mexico Emissions Data [<a href="https://www.epa.gov/air-emissions-modeling/2014-2016-version-7-air-emissions-modeling-platforms">National Emissions Inventory Collaborative, 2016v1 emissions modeling platform</a>] <ul> <li>Mexico_2016_point_interpolated_02mar2018_v0.csv</li> </ul> </li> <li> North American Regional Reanalysis Meteorological Data [<a href="https://psl.noaa.gov/data/gridded/data.narr.monolevel.html">NOAA</a>] <ul> <li>rhum.2m.mon.mean.nc</li> <li>uwnd.10m.mon.mean.nc</li> <li>vwnd.10m.mon.mean.nc</li> </ul> </li> <li> Cigarette smoking data [<a href="https://doi.org/10.1186/1478-7954-12-5">Dwyer-Lindgren et al. (2014)</a>] <ul> <li>smokedatwithfips_1996-2012.csv</li> </ul> </li> <li> Synthetic pediatric asthma data [<em>Note:<strong> synthetic data!</strong> Simulated to match the format, but not the observations, from the <a href="https://www.dshs.texas.gov/texas-health-care-information-collection">Texas Health Care Information Collection (THCIC), Texas DSHS</a></em>] <ul> <li>synth-ped-asthma-data.csv</li> </ul> </li> <li> Texas state shape file [<a href="https://www.census.gov/geographies/mapping-files/time-series/geo/carto-boundary-file.html">US Census</a>] <ul> <li>texas-state-sf.RDS</li> </ul> </li> <li> US ZIPcode-to-county data crosswalk [<a href="https://mcdc.missouri.edu/applications/geocorr2014.html">Missouri Census Data Center</a>] <ul> <li>tx-zip-to-county.csv</li> </ul> </li> </ol> <p>Code and supplementary material from this analysis, as well as more detailed data descriptions, are available at: <a href="https://github.com/nbwikle/estimating-interference">https://github.com/nbwikle/estimating-interference</a></p>
Data for the publication "Incorporation of inline warm rain diagnostics into the COSP2 satellite simulator for process-oriented model evaluation"
<p>Michibata et al. (2019), currently under peer-review for publication in <em>Geoscientific Model Development</em>, incorporated a diagnostic tool for warm rain microphysics into the CFMIP Observation Simulator Package (COSP; Bodas-Salcedo et al. 2011; Swales et al., 2018), designed to evaluate model representations of aerosol–cloud–precipitation interactions at a fundamental process-level. The tool automatically generates two diagnostics related to warm rain microphysics during COSP execution in a host model. One is the contoured frequency by optical depth diagram (CFODD), which visualizes a cloud-to-rain microphysical vertical structure (Suzuki et al., 2015). The other diagnostic is a global map of warm rain fraction classified as non-precipitating clouds (< –15 dBZ<sub>e</sub>), drizzling clouds (–15 < dBZ<sub>e</sub>< 0), and precipitating clouds (0 < dBZ<sub>e</sub>).</p> <p>This repository contains the MIROC6/COSP2 input data and A-Train satellite statistics used in Michibata et al. (2019). A sample of the post-processing scripts for visualization using the GrADS software is also included in this repository.</p>
Replication Package for the paper "Conversing with business process-aware Large Language Models: the BPLLM framework"
<p>Replication Package for the research paper "<em>Conversing with business process-aware Large Language Models: the BPLLM framework</em>".</p> <p>The package includes the process models, the questions (and expected answers), the results of the qualitative evaluation, and the Hugging Face links to the fine-tuned versions of Llama 3.1 8B employed in the quantitative evaluation of the framework.</p> <p>In particular, the process models are:</p> <ul> <li>The natural language Directly-follows graph (DFG) of the Food Delivery process: <em>food_delivery_activities.txt</em> for the definition of the activities and <em>food_delivery_flow.txt</em> for the sequence flow.</li> <li>The BPMN model of the Food Delivery, E-commerce, and Reimbursement processes: <em>ecommerce.bpmn</em>, <em>food_delivery.bpmn</em>, and <em>reimbursement.bpmn</em>.</li> </ul> <p>The datasets with the questions and the expected answers are:</p> <ul> <li><em>1_questions_answers_not_refined_for_DFG.csv</em> ;</li> <li><em>1.1_questions_answers_refined_for_DFG.csv</em> ;</li> <li><em>2_questions_answers_not_refined.csv</em> ;</li> <li><em>3_questions_answers_refined.csv</em> ;</li> <li><em>4_questions_answers_different_processes.csv</em> ;</li> <li><em>5_questions_answers_similar_processes.csv</em> ;</li> <li><em>6_questions_answers_refined_ft.csv</em> .</li> </ul> <p>The complete results of the qualitative evaluation are contained in the file <em>qualitative_experiments_results.pdf</em>.</p> <p>The Hugging Face links to the fine-tuned versions of Llama 3.1 8B are reported in <em>hf_links_finetuned_models.pdf</em>.</p>
Global topsoil SOC stock from 1981 to 2018 estimated by combining process-based model and space-for-time digital soil mapping
<p>This dataset include the topsoil (0-30cm) soil organic carbon (SOC) stocks in mineral soils under major land classes (forest, grassland, shrub land, savannas, cropland, cropland/natural vegetation mosaic, and sparely vegetated land) from 1981 to 2018. The long-time series of SOC stocks were estimated by using a space-for-time digital soil mapping (DSMst) model where the RothC-simulated SOC stocks were incorporated as one of the dynamic covariates of the DSMst model.</p> <p>The detail information on the products were given below:</p> <p>Name: DSMst-RothC 5-km global topsoil SOC stock products</p> <p>Period: 1981-2018</p> <p>Spatial resolution: 0.041666667 degree</p> <p>Temporal resolution: 1 year</p> <p>CRS: geographic latitude/longitude (EPSG:4326 - WGS 84 – Geographic)</p> <p>Extent: -180°, -90°: 180°, 90°</p> <p>Data format: GeoTIFF</p> <p>Compression: LZW</p> <p>Data type: Float32</p> <p>Unit: t C ha<sup>-1</sup></p>
Processed Synthetic Real-World Data for tristate modelling
<p>This model learning dataset is created out of the <a href="https://zenodo.org/record/7409763">Raw Synthetic RWD</a> raw dataset, including some of the original attributes. It is distributed in JOBLIB files, where .joblib files contain the vectors and _ids.joblib contain the ID of the person from which each vector is extracted.</p> <p>This is useful in case it is needed to map the vectors to metadata about the people that are found in the original raw dataset. Note that corresponds to , or , depending on the dataset.</p> <p>The split is roughly 60% of the people are in the training dataset, and 20% in each of the validation and the testing datasets. The input attributes are the age, the short-term averages and the trends of the current week’s BMI, steps walked, calories burned, sleep quality, mood and water consumption, as well as the previous week’s short-term average and trend of the answer to the health self-assessment question.</p> <p>The outcome to be predicted is a tristate quantized version of the health self-assessment answer to be given in the current week. The dataset is normalized based on the training set. The means and standard deviations used can be found in the train_statistics.joblib file. Finally, the output_descriptions.joblib file contains descriptions of the outcomes to be predicted (not actually needed, since included here).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.