Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
76,402,788
datasets available to search
ShareScore release 0.7.1
Dataset results
76,402,788 results
Data set for the journal article ''Nanoscale chemical reaction exploration with a quantum magnifying glass''
<div>This data set includes the raw data of the esterification and hydrogenation discussed in the journal article alongside with the Scine Puffin Singularity container, steering protocol files, Swoose parameters, (pre-)releases of the software, and Python scripts for individual steps without the graphical user interface to reproduce the data.</div>
BLASTNet Simulation Dataset
<p>Go to <a href="https://blastnet.github.io/">https://blastnet.github.io/</a> to access and download this reacting and non-reacting flow physics simulations.</p> <p><strong>Mission</strong></p> <p>BLASTNet 2.1 was developed to provide the researchers in reacting and non-reacting flow physics communities with high-fidelity simulation datasets in a convenient format for ML applications. With ~5 TB, 765 full-domain samples, and 36 configurations, BLASTNet can effectively address these gaps and aid in fostering open/fair ML development within reacting and non-reacting flow physics communities.</p> <p><strong>Application</strong></p> <p>This data is useful for fluid flows in a wide range of ML applications tied to automotive, propulsion, energy, and the environment. Specifically, scientific engineering tasks related to these domains may include turbulent closure modeling, spatio-temporal modeling, and inverse modeling.<br> </p>
Database of fitted spectra for: Changing-Look AGNs - I. Tracking the transition on the main sequence of quasars
<h3>Results from the spectral fitting for a sample of changing-look active galactic nuclei (AGNs) with SDSS spectroscopy using PyQSOFit.</h3>
Generative convective parametrization of a dry atmospheric boundary layer
<p>The repository contains simulation snapshots of a dry convective boundary layer (CBL). The snapshots comprise horizontal snapshots of vertical velocity (w) and buoyancy (b) field at three heights, namely z/h(t) = 0.2, 0.5, 1.0. Further, the Python scripts for the Generative Adversarial Network (GAN) are also provided, as well as the DNS renormalization procedure.</p>
GlobalHighPM₂.₅: Global Daily Seamless 1 km Ground-Level PM₂.₅ Dataset over Land (2017–Present)
<p>GlobalHighPM<sub>2.5</sub> is part of a series of long-term, seamless, global, high-resolution, and high-quality datasets of air pollutants over land (i.e., GlobalHighAirPollutants, GHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>This dataset contains input data, analysis codes, and generated dataset used for the following article. If you use the GlobalHighPM<sub>2.5</sub> dataset in your scientific research, please cite the following reference (Wei et al., NC, 2023):</p> <ul> <li> <p>Wei, J., Li, Z., Lyapustin, A., Wang, J., Dubovik, O., Schwartz, J., Sun, L., Li, C., Liu, S., and Zhu, T. <a href="https://weijing-rs.github.io/publications/Wei_et_al-NC-2023.pdf" target="_blank" rel="noopener">First close insight into global daily gapless 1 km PM<sub>2.5</sub> pollution, variability, and health impact</a>. <em>Nature Communications</em>, 2023, 14, 8349. https://doi.org/10.1038/s41467-023-43862-3</p> </li> </ul> <p><strong>Input Data</strong></p> <p>Relevant raw data for each figure (compiled into a single sheet within an Excel document) in the manuscript.</p> <p><strong>Code</strong></p> <p>Relevant Python scripts for replicating and ploting the analysis results in the manuscript, as well as codes for converting data formats.</p> <p><strong>Generated Dataset</strong></p> <p>Here is the first big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) global ground-level PM<sub>2.5</sub> dataset over land from 2017 to the present. This dataset exhibits high quality, with cross-validation coefficients of determination (CV-R<sup>2</sup>) of 0.91, 0.97, and 0.98, and root-mean-square errors (RMSEs) of 9.20, 4.15, and 2.77 µg m<sup>-3</sup> on the daily, monthly, and annual bases, respectively.</p> <p><strong>Due to data volume limitations, </strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2022 </strong>is accessible at: <strong><a href="../records/10795661">GlobalHighPM2.5 (2022)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2021 </strong>is accessible at: <strong><a href="../records/10398385">GlobalHighPM2.5 (2021)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2020 </strong>is accessible at: <strong><a href="../records/10402639">GlobalHighPM2.5 (2020)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2019 </strong>is accessible at: <strong><a href="../records/10402723">GlobalHighPM2.5 (2019)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2018 </strong>is accessible at: <strong><a href="../records/10402824">GlobalHighPM2.5 (2018)</a></strong></p> <p> all (including <strong>daily</strong>) data for the year <strong>2017 </strong>is accessible at: <strong><a href="../records/10403497">GlobalHighPM2.5 (2017)</a></strong></p> <p> continuously updated...</p> <p><strong>More GHAP datasets for different air pollutants are available at: <a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>
Dataset - Decrypting lysine deacetylase inhibitor action and protein modifications by dose-resolved proteomics
<h4><strong>Dataset Summary</strong></h4> <p>Lysine deacetylase inhibitors (KDACis) are approved for cutaneous T-cell lymphoma (CTCL), peripheral T-cell lymphoma (PTCL), and multiple myeloma. Despite the mechanism of action(s) (MoA) remains elusive, these inhibitors lead to increasing acetylation levels of histones and other proteins, altered gene expression and cell death. To characterize the MoA of these drugs in more detail, we systematically measured dose-dependent changes in protein expression, acetylation, and phosphorylation in response to 21 clinical and pre-clinical KDACis. MV4-11 cells were treated for 6 h with 1 vehicle control and 10 increasing doses of the respective drug (from 100 pM to 30 mM). Proteins were digested with trypsin, and the resulting 11 peptide preparations corresponding to one drug dose each were encoded by stable isotopes (tandem mass tags, TMT-11plex) and combined. Acetylated peptides were subsequently enriched by immunoprecipitation and phosphopeptides by immobilized metal affinity chromatography (IMAC). PTM-carrying and unmodified peptides were analyzed separately by liquid chromatography tandem mass spectrometry (LC-MS/MS) for peptide and protein identification and quantification. Additionally, Vorinostat and Panobinostat were also recorded as time-dependent experiments at their pEC50 concentration, respectively. </p> <h4><strong>Dataset structure</strong></h4> <p>Here, we provide all curve data processed with CurveCurator v0.4.0 (<a href="https://github.com/kusterlab/curve_curator">https://github.com/kusterlab/curve_curator</a>). Each drug is a zip folder containing acetylome, phosphoproteome, and fullproteome data. Next to each data set is the toml parameter file used to generate the curves.txt and dashboard.html files. Time-dependent data is indicated by "td" and dose-dependent data is indicated by "dd".</p> <p> </p>
Cebulka (Polish dark web cryptomarket and image board) messages data
<h3><strong>General Information</strong></h3> <p>1. <strong>Title of Dataset</strong></p> <p>Cebulka (Polish dark web cryptomarket and image board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Patrycja Cheba (Jagiellonian University); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: “Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade” (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4.<strong> Data Source</strong></p> <p>Polish dark web cryptomarket and image board called Cebulka (<a href="http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php">http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php</a>). </p> <p>5. <strong>Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project focuses on studying internet behavior concerning disruptive actions, particularly emphasizing the online narcotics market in Poland. The research seeks to (1) investigate how the open internet, including social media, is used in the drug trade; (2) outline the significance of darknet platforms in the distribution of drugs; and (3) explore the complex exchange of content related to the drug trade between the surface web and the darknet, along with understanding meanings constructed within the drug subculture.</p> <p>Within this context, Cebulka is identified as a critical digital venue in Poland’s dark web illicit substances scene. Besides serving as a marketplace, it plays a crucial role in shaping the narratives and discussions prevalent in the drug subculture. The dataset has proved to be a valuable tool for performing the analyses needed to achieve the project’s objectives.</p> <h3><strong>Data Content</strong></h3> <p>6. <strong>Data Description</strong></p> <p>The data was collected in three periods, i.e., in January 2023, June 2023, and January 2024.</p> <p>The dataset comprises a sample of messages posted on Cebulka from its inception until January 2024 (including all the messages with drug advertisements). These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. The dataset is organized into two directories. The “cebulka_adverts” directory contains posts related to drug advertisements (both advertisements and comments). In contrast, the “cebulka_community” directory holds a sample of posts from other parts of the cryptomarket, i.e., those not related directly to trading drugs but rather focusing on discussing illicit substances. The dataset consists of 16,842 posts.</p> <p>7. <strong>Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>8. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following files:</p> <ul> <li>Zipped .txt files (“cebulka_adverts.zip” and “cebulka_community.zip”) containing all messages. These files are organized into individual directories that mirror the folder structure found on Cebulka.</li> <li>Two .csv files that list all the messages, including file names and the content of each post. The first .csv lists messages from “cebulka_adverts.zip,” and the second .csv lists messages from “cebulka_community.zip.”</li> </ul> <h3><strong>Ethical Considerations</strong></h3> <p>9. <strong>Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the messages, utilizing automated systems for irreversible hashing. Recognizing that automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>
Hyperreal Talk (Polish clear web message board) messages data
<h3><strong>General Information</strong></h3> <p>1.<strong> Title of Dataset</strong></p> <p>Hyperreal Talk (Polish clear web message board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: “Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade” (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4.<strong> Data Source</strong></p> <p>Polish clear web message board called Hyperreal Talk (<a href="https://hyperreal.info/talk/">https://hyperreal.info/talk/</a>).</p> <p>5.<strong> Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project delves into internet dynamics within disruptive activities, specifically focusing on the online drug trade in Poland. It aims to (1) examine the utilization of the open internet, including social media, in the drug trade; (2) delineate the role of darknet environments in narcotics distribution; and (3) uncover the intricate flow of drug trade-related content and its meanings between the open web and the darknet, and how these meanings are shaped within the so-called drug subculture.</p> <p>The Hyperreal Talk forum emerges as a pivotal online space on the Polish internet, serving as a hub for discussions and the exchange of knowledge and experiences concerning drug use. It plays a crucial role in investigating the narratives and discourses that shape the drug subculture and the broader societal perceptions of drug consumption. The dataset has been instrumental in conducting analyses pertinent to the earlier project goals.</p> <p>6. <strong>Collection Method</strong></p> <p>The dataset was compiled using the Scrapy framework, a web crawling and scraping library for Python. This tool facilitated systematic content extraction from the targeted message board.</p> <p>7.<strong> Collection Date</strong></p> <p>The data was collected in two periods, i.e., in September 2023 and November 2023.</p> <h3><strong>Data Content</strong></h3> <p>8. <strong>Data Description</strong></p> <p>The dataset comprises all messages posted on the Polish-language Hyperreal Talk message board from its inception until November 2023. These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. The dataset is organized into two directories: “hyperreal” and “hyperreal_hidden.” The “hyperreal” directory contains accessible posts without needing to log in to Hyperreal Talk, while the “hyperreal_hidden” directory holds posts that can only be viewed by logged-in users. For each directory, a .txt file has been prepared detailing the structure of the message board folders from which the posts were extracted. The dataset includes 6,248,842 posts.</p> <p>9.<strong> Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>10. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following files:</p> <ul> <li>Zipped .txt files (hyperreal.zip) containing messages that are visible without logging into Hyperreal Talk. These files are organized into individual directories that mirror the folder structure found on the Hyperreal Talk message board.</li> <li>Zipped .txt files (hyperreal_hidden.zip) containing messages that are visible only after logging into Hyperreal Talk. Similar to the first type, these files are organized into directories corresponding to the website’s folder structure.</li> <li>A .csv file that lists all the messages, including file names and the content of each post.</li> </ul> <h3><strong>Accessibility and Usage</strong></h3> <p>11.<strong> Access Conditions</strong></p> <p>The data can be accessed without any restrictions.</p> <p>12. <strong>Related Documentation</strong></p> <p>Attached are .txt files detailing the tree of folders for “hyperreal.zip” and “hyperreal_hidden.zip.”</p> <p>Documentation on the Python regular expressions used for scraping, cleaning, processing, and anonymizing the data can be found on GitHub at the following URLs:</p> <ul> <li><a href="https://github.com/LeszekSwieca/Project_2021-43-B-HS6-00710">https://github.com/LeszekSwieca/Project_2021-43-B-HS6-00710</a></li> <li><a href="https://github.com/HaitaoShi/Scrapy_hyperreal">https://github.com/HaitaoShi/Scrapy_hyperreal</a>"</li> </ul> <h3><strong>Ethical Considerations</strong></h3> <p>13. <strong>Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the messages, utilizing automated systems for irreversible hashing. Recognizing that scraping and automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>
Dopek.eu (Polish clear web and dark web message board) messages data
<h3><strong>General Information</strong></h3> <p>1. <strong>Title of Dataset</strong></p> <p>Dopek.eu (Polish clear web and dark web message board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: “Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade” (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4. <strong>Data Source</strong></p> <p>Clear web and dark web message board called dopek.eu (<a href="https://dopek.eu/">https://dopek.eu/</a>). </p> <p>5.<strong> Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project delves into internet dynamics within disruptive activities, specifically focusing on the online drug trade in Poland. It aims to (1) examine the utilization of the open internet, including social media, in the drug trade; (2) delineate the role of darknet environments in narcotics distribution; and (3) uncover the intricate flow of drug trade-related content and its meanings between the open web and the darknet, and how these meanings are shaped within the so-called drug subculture.</p> <p>The dopek.eu forum emerges as a pivotal online space on the Polish internet, serving as a hub for trading, discussions, and the exchange of knowledge and experiences concerning the use of the so-called new psychoactive substances (designer drugs). The dataset has been instrumental in conducting analyses pertinent to the earlier project goals.</p> <p>6.<strong> Collection Method</strong></p> <p>The dataset was compiled using the Scrapy framework, a web crawling and scraping library for Python. This tool facilitated systematic content extraction from the targeted message board.</p> <p>7.<strong> Collection Date</strong></p> <p>The data was collected in October 2023.</p> <h3><strong>Data Content</strong></h3> <p>8.<strong> Data Description</strong></p> <p>The dataset comprises all messages posted on dopek.eu from its inception until October 2023. These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. A .txt file has been prepared detailing the structure of the message board folders from which the posts were extracted. The dataset includes 171,121 posts.</p> <p>9.<strong> Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>10. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following types of files:</p> <ul> <li>Zipped .txt files (dopek.zip) containing all messages (posts).</li> <li>A .csv file that lists all the messages, including file names and the content of each post.</li> </ul> <h3><strong>Accessibility and Usage</strong></h3> <p><strong>11. Access Conditions</strong></p> <p>The data can be accessed without any restrictions.</p> <p><strong>12. Related Documentation</strong></p> <p>Attached are .txt files detailing the tree of folders for “dopek.zip”.</p> <h3><strong>Ethical Considerations</strong></h3> <p><strong>13. Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the posts, utilizing automated systems for irreversible hashing. Recognizing that scraping and automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>
Dataset of "Fast carbon dioxide–epoxide cycloaddition catalyzed by metal and metal-free ionic liquids for designing non-isocyanate polyurethanes"
<p>The recycling of industrially produced greenhouse gases, such as CO2, into high-value-added chemicals is one of the most relevant strategies for reaching climate targets. A two-step strategy for designing non-isocyanate polyurethanes (NIPUs) from renewable carbon dioxide (CO2) using environmentally friendly conditions and catalysts is investigated. The first reaction step efficiently converts a mono-epoxidized monomer (phenyl glycidyl ether) into cyclic carbonates under mild reaction conditions and supercritical CO2, using imidazolium ionic liquids (ILs) as catalysts into cyclic carbonates. The DFT calculations suggested a comprehensive mechanistic pathway for the IL-catalyzed CO2-epoxy reaction showing a rate-determining step of the initial epoxide ring opening and the direct participation of IL-anions.</p>
Dataset for Accuracy of Grid-Connected Photovoltaic Power Plant: A Novel Approach Using Hybrid Variational Mode Decomposition and CNN-LSTM Model
<p>This research paper introduces a deep learning hybrid model employing Convolutional Neural Network Long Short-Term Memory (CNN-LSTM) for short-term photovoltaic (PV) solar energy forecasting.The proposed method integrates the Variational Mode Decomposition (VMD) algo-rithm with the CNN-LSTM model to predict PV power generation from a solar farm in Boussada, Algeria, from January 1, 2019, to December 31, 2020. The performance of the developed model is benchmarked against other deep learning models (VMD-CNN, VMD-LSTM, CNN-LSTM) across various time horizons (15, 30, and 60 minutes) to provide a comprehensive evaluation. Our findings exhibit greater performance of the developed model compared to other architectures, showcasing promising results in solar power forecasting. This research contributes to the main goal of enhancing EMS by providing accurate solar energy forecasts.</p>
EOSC Task Force on FAIR Metrics and Data Quality: FAIR Evaluation community survey 2023
<p>The EOSC-A FAIR Metrics and Data Quality Task Force (TF) supported the European Open Science Cloud Association (EOSC-A) by providing strategic directions on FAIRness (Findable, Accessible, Interoperable, and Reusable) and data quality. The Task Force conducted a survey using the <a href="https://ec.europa.eu/eusurvey/">EUsurvey tool</a> between 15.11.2022 and 18.01.2023, targeting both developers and users of FAIR assessment tools. The survey aimed at supporting the harmonisation of FAIR assessments, in terms of what it evaluated and how, across existing (and future) tools and services, as well as explore if and how a community-driven governance on these FAIR assessments would look like. The survey received 78 responses, mainly from academia, representing various domains and organisational roles. This is the anonymised survey dataset in csv format; most open-ended answers have been dropped. The codebook contains variable names, labels, and frequencies.</p>
Belvedere Glacier long-term monitoring Open Data
<p><strong>Introduction </strong></p> <p>This dataset contains extensive, long-term monitoring data on the Belvedere Glacier, a debris-covered glacier located on the east face of Monte Rosa in the Anzasca Valley of the Italian Alps. The data is derived from photogrammetric 3D reconstruction of the full Belvedere Glacier and includes:</p> <ul> <li><strong>dense point clouds</strong> obtained with UAV-based MVS covering the entire glacier body</li> <li>high-resolution<strong> </strong><strong>orthophotos</strong></li> <li>high-resolution<strong> </strong><strong>DEMs</strong></li> </ul> <p>Since 2015, in-situ survey of the glacier have been conducted annually using fixed-wing UAVs until 2020 and quadcopters from 2021 to 2022 to remotely sense the glacier and build high-resolution photogrammetric models. A set of ground control points (GCPs) were materialized all over the glacier area, both inside the glacier and along the moraines, and surveyed (nearly-) yearly with topographic-grade GNSS receivers (Ioli et al., 2022).</p> <p>For the period from 1977 to 2001, historical analog images, digitalized with photogrammetric scanners and acquired from aerial platforms, were used in combination with GCPs obtained from recent photogrammetric models (De Gaetani et al., 2021).</p> <p>Before downloading them, you can explore the photogrammetric point clouds of the Belvedere Glacier within web app based on Potree from <a href="https://thebelvedereglacier.it/" target="_blank" rel="noopener">https://thebelvedereglacier.it/</a> (use a web browser from a desktop/laptop for the best experience). Additionally, from here you can also visualize and download the coordinates of the GCPs measured by GNSS every year since 2015.</p> <p> </p> <p><strong>Belvedere Glacier </strong></p> <p>The Belvedere Glacier is an important temperate alpine glacier located on the east face of Monte Rosa in the Anzasca Valley of Italy. The Belvedere Glacier is of particular importance among alpine glaciers because it is a debris-covered glacier and it reaches its lowest elevation at about 1800 m a.s.l. Over the last century, the Belvedere Glacier has experienced extraordinary dynamics, such as a surge-like movement or the formation of a supraglacial lake, which seriously threatened the nearby community of Macugnaga.</p> <p> </p> <p><strong>Data organization</strong></p> <p>The data are organized by year in compressed zip folders named <em>belvedere_YYYY.zip</em>, which can be downloaded independently. Each folder contains all data available for that year (i.e. photogrammetric point clouds, orthophotos, and DEMs) and the corresponding metadata. Metadata is provided as a .json file which contains all the main information for data usage. Point clouds are saved in compressed las format (<em>.laz</em>)<em> </em>and they can be inspected e.g., with CloudCompare. Orthophotos and DEMs are georeferenced images (<em>.tif</em>) that can be inspected with any GIS software (e.g., <em>QGIS</em>).</p> <p>Large point clouds are subdivided into regular tiles, which are numbered in a progressive row-wise order from the bottom-left corner of the point cloud bounding box.</p> <p>All the files are named according to the following naming schema:</p> <p>"belv_YYYY_surveyplatform_datatype[_resolution][vertical_datum][-tile_number].extension"</p> <p>where: </p> <ul> <li>YYYY: is the year of the survey</li> <li>surveyplatform: can be either "uav" for the UAV-based photogrammetry survey or "histo" for the historical aerial datasets.</li> <li>datatype: can be either "pcd" for point clouds, "orthophoto" for orthophotos and "dsm" for DSMs. </li> <li>resolution: on-ground resolution of each pixel in meters. This applies only to raster data (orthophoto and DSMs)</li> <li>vertical_datum: if the DSM is given in orthometric coordinates, the label "ortho" is present in the filename, otherwise the height of the dataset is supposed to be ellipsoidal.</li> <li>tile: tile number, if the data is tiled to avoid large files.</li> </ul> <p><strong>Data Usage</strong></p> <p>This dataset can be used to estimate glacier velocities, volume variations, study geomorphological processes such as the process of moraine collapse, or derive other information on glacier dynamics. If you have any requests on the data provided, data acquisition, or the raw data themselves, you are encouraged to contact us.</p> <p> </p> <p><strong>Contributions</strong></p> <p>The monitoring activity carried out on the Belvedere Glacier was designed and conducted jointly by the Department of Civil and Environmental Engineering (DICA) of Politecnico di Milano and the Department of Environment, Land and Infrastructure Engineering (DIATI) of Politecnico di Torino. The DREAM projects (DRone tEchnnology for wAter resources and hydrologic hazard Monitoring), involving teachers and students from Alta Scuola Politecnica (ASP) of Politecnico di Torino and Milano, contributed to the campaign from 2015 to 2017.</p> <p> </p> <p><strong>Acknowledgements</strong></p> <div>The authors thank CGR SpA for digitizing the historical images (1977, 1991, 2001, 2009) and making them available to the authors for the photogrammetric processing.</div> <div>The authors thank all students and collaborators contributing to the Alta Scuola Politecnica projects DREAM 1, DREAM 2, and DREAM 3 (DRone tEchnnology for wAter resources and hydrologic hazard Monitoring). </div> <div> </div> <div> </div> <p><strong>If you use the data, please, cite these our pubblications:</strong></p> <p>Ioli, F., Dematteis, N., Giordan, D., Nex, F., Pinto, L., Deep Learning Low-cost Photogrammetry for 4D Short-term Glacier Dynamics Monitoring. <em>PFG</em> (2024). <a href="https://doi.org/10.1007/s41064-023-00272-w" target="_blank" rel="noopener">https://doi.org/10.1007/s41064-023-00272-w</a></p> <p>Ioli, F.; Bianchi, A.; Cina, A.; De Michele, C.; Maschio, P.; Passoni, D.; Pinto, L. Mid-Term Monitoring of Glacier’s Variations with UAVs: The Example of the Belvedere Glacier. Remote Sensing, 14, 28 (2022). <a href="https://doi.org/10.3390/rs14010028" target="_blank" rel="noopener">https://doi.org/10.3390/rs14010028</a></p> <p>De Gaetani, C.I.; Ioli, F.; Pinto, L. Aerial and UAV Images for Photogrammetric Analysis of Belvedere Glacier Evolution in the Period 1977–2019. Remote Sensing, 13, 3787 (2021). <a href="https://doi.org/10.3390/rs13183787" target="_blank" rel="noopener">https://doi.org/10.3390/rs13183787</a></p>
Energy Cycle Characteristics for 5G/6G Networks Supported by RES, UAVs, and RISs
<h2><strong>Overview</strong></h2> <p>The following dataset presents the energy cycle characteristics for 5G/6G mobile systems supported by Renewable Energy Sources (RES) and/or Unmanned Aerial Vehicles (UAVs) and Reconfigurable Intelligent Surfaces (RISs). In addition, within the dataset, the energy gain related to the engagement of RES within the Radio Access Network (RAN) has also been distinguished.</p> <h2><strong>Scenario</strong></h2> <p>The considered network scenario includes 8 three- (<em>_results_gcas.csv</em>) or one-cell (<em>_results_scas.csv</em> & <em>_results_kras.csv</em>) base stations (BSs) placed within the Poznan city (surroundings of the old market) and supported by Renewable Energy Sources — photovoltaic panels (PVs) and/or wind turbines (WTs). The aforementioned base stations can be treated as stationary towers or mobile access points (e.g., drones/UAVs). Those latter have been additionally equipped with RIS devices, which are able to reflect and manipulate a radio signal to influence occurrences such as interferences, coverage, or human exposure. However, the use of RISs has been taken into account only to evaluate the impact of the engagement of such devices on the energy side of the mobile system, omitting the changes in radio characteristics. The network traffic has been assumed to be fixed (64 mobile users (UEs) with 100 Mbps downlink — DL, and 25 Mbps uplink — UL, per each), however, its density in specific parts of the city is modeled randomly for each simulation run. The simulation runs have been performed for 4 dates (vernal equinox, summer solstice, autumn equinox, winter solstice), each one from a different season of the year. The aim of such an approach was to highlight the impact of the time of the day and the year on the energy gain obtained thanks to enabling RES generators. The weather conditions assumed within the simulation are typical for the climate in Poland. </p> <h2><strong>Methodology</strong></h2> <p>The energy-cycle calculations (system's power consumption, renewable energy production, and excessive energy storage) have been based on the mathematical formulas from the scientific literature and performed within the digital simulation runs by using the Green Radio Access Network Design (GRAND) tool (developed by teams from the Ghent University & Poznan University of Technology). The UE-BS association process within the mobile system has been done by doing multi-objective optimization using the Gurobi software, which has taken into account parameters like path loss, predicted power consumption of BSs, and guaranteed DL & UL bit rates for UEs.</p> <h2><strong>Simulation setup</strong></h2> <p>The setup of the input parameters for used mathematical models (power consumption, energy generation, energy storage) has been done in accordance with the values attached within the delivered literature positions (cited within the publications included in the <em>Related works</em> section of the following dataset) and adjusted to the considered study. Furthermore, the data used to model the network environment (building distribution, coverage area, base stations' locations) as well as to predict weather conditions are the real data (for the year 2022) collected by the city hall of Poznan, one of the Polish mobile operators, and weather stations placed in Poznan, respectively. The number of simulation runs performed has been equal to 10 (each run has included energy-cycle calculations for 4 seasons of the year), with the time step of a single run set to 1 hour of the day.</p> <h2><strong>Results</strong></h2> <p>The results of the aforementioned investigations have been included in the attached files, which can be described as follows:</p> <h3><strong>File <em>_results_gcas.csv</em></strong></h3> <p>The first column denotes the date (season of the year), for which the values have been obtained. The columns from second to fifth present observed values of the State of Charge (SoC) of a battery system (in %) for a single network cell on average in a time step. Those columns are the obtained values for the RAN, in which no RES, only PVs, only WTs, and both types of RES generators have been enabled, respectively. </p> <h3><strong>Files <em>_results_scas.csv</em> & <em>_results_kras.csv</em></strong></h3> <p>The first column denotes the date (season of the year), for which the values have been obtained. The second and third columns denote the number of drone base station (DBS) exchanges within the wireless system on average in a particular time step, where no RES and only PVs are enabled, respectively. The fourth and fifth columns present the conventional (fossil-fuels-based) energy consumption (in kWh) for the whole system in a specific time step, in which no RES and only PVs are engaged for all the access nodes. The sixth column is the energy savings (in kWh) related to the use of RES generators within the mobile network. Furthermore, the seventh and eighth columns represent the amount of renewable energy harvested from the solar radiation in total and the peak value of this amount observed during the entire day, respectively.</p> <h2><strong>Acknowledgment</strong></h2> <p>More details about the conducted studies have been described within the attached papers (<em>Related works</em> section). The data has been collected within the COST CA10210 INTERACT. M. Deruyck is a Post-Doctoral Fellow of the FWO-V (Research Foundation – Flanders, ref: 12Z5621N). The work (including the following dataset preparation) by A. Samorzewski and A. Kliks was realized within project no. 2021/43/B/ST7/01365 funded by the National Science Center in Poland.</p>
The local extinction of Cedrus atlantica in the Iberian Peninsula could have been completed due to biological interaction
<p>This data set is used to explore the possibility that <em>Cedrus atlantica</em> (Endl.) Carrière and <em>Pinus nigra</em> Arnold could have interacted in the past, mutually excluding each other in the areas with suitable conditions for both species and, where, ultimately, the one that was most competitive would remain. The species show very well differenciated niches and a distribution of their habitats segregated by continents (<em>P. nigra</em> in Europe and <em>C. atlantica</em> in Africa), which responds to differences in climatic affinities. However, the contact of their distributions in bordering areas suggests that <em>C. atlantica</em> maintained its presence in the Iberian Peninsula until recent times, and that <em>P. nigra</em> could have displaced it due to its higher prevalence on the continent.</p>
Identification of an altitudinal migration pattern of Abies pinsapo in the Baetic Mountains through the presence of its life stages
<p>This data set is used to explore the altitudinal shift of <em>Abies pinsapo</em> Boiss. in the Baetic System. We analysed the potential distribution of the realised and reproductive niches of <em>A. pinsapo</em> populations in the Ronda Mountains (Southern Spain) by using species distribution models (SDMs) for two life stages within the current populations. The realised and reproductive niches of <em>A. pinsapo</em> are different to one another, which may indicate a displacement in its altitudinal distribution.</p>
nuts-STeauRY dataset: hydrochemical and catchment characteristics dataset for large sample studies of Carbon, Nitrogen, Phosphorus and Silicon in french watercourses
<p><strong>nuts-STeauRY dataset: hydrochemical and catchment characteristics dataset for large sample studies of Carbon, Nitrogen, Phosphorus and Silicon in French watercourses</strong></p> <p>Antoine Casquin, Marie Silvestre, Vincent Thieu</p> <p>10.5281/zenodo.10830852</p> <p>v0.1, 18<sup>th</sup> March 2024</p> <p><strong>Brief overview of data: </strong></p> <p>· Carbon and nutrients data for 5470 continental French catchments</p> <p>· Modelled discharge for 5128 of catchments out of 5470</p> <p>· Geopackages with catchment delineations and outlets</p> <p>· DEM conditioned to delimit additional catchments</p> <p>· Land-use and climatic data for 5470 continental French catchments</p> <p><strong>Citation of this work<br></strong></p> <p>A data paper is currently being submitted with details of methods and results. Once published, it will be the preferential source to cite. The data paper will be link to the new version of the dataset that will be updated on doi.org/10.5281/zenodo.10830852. If you use this dataset in your research or report, you must cite it.</p> <p><strong>Motivations</strong></p> <p>Data was collected and curated for the nuts-STeauRY project (<a href="http://nuts-steaury.cnrs.fr">http://nuts-steaury.cnrs.fr</a>), which deployed a national generic land to sea modelling chain.</p> <p>Data was primarily used (see related works):</p> <ol> <li>To calibrate concentrations of dissolved organic carbon and dissolve silica in headwaters</li> <li>To validate spatially and temporally the modelling chain (DOC, NO3-, NH4+, TP, SRP, DSi)</li> </ol> <p>Hydrochemical large sample datasets have numerous other uses: trends computations elucidate transfer mechanisms, machine learning, retrospective studies etc.</p> <p>The objective here is to provide a large sample curated dataset of carbon and nutrients concentrations along with modelled discharges, catchment characteristics and delimitations for the continental France. Such large sample dataset aims at easing the large sample studies over France and/or Europe. Although part of the data gathered here is obtainable via public sources, the catchments delineations, their characteristics and modelled hydrology were note not publicly available yet. Moreover, a unification of units and detection and removal of outliers was performed on carbon and nutrients data.</p> <p><strong>Data sources & processing</strong></p> <p>Sampling points where snapped on the CCM database v2.1 (<a href="http://data.europa.eu/89h/fe1878e8-7541-4c66-8453-afdae7469221">http://data.europa.eu/89h/fe1878e8-7541-4c66-8453-afdae7469221</a>)(Vogt et al., 2007) and catchments were delineated using a 100m resolution Digital Elevation Model (DEM) conditioned by the hydrographic network and elementary catchments’ delineations of the CCM data v2.1. <strong>More than 6000 catchments were delineated and screened manually</strong> to check consistency: 5470 were retained<strong>.</strong></p> <p>Nutrient data was collected mainly through the Naiades portal (<a href="https://naiades.eaufrance.fr/">https://naiades.eaufrance.fr/</a>), a database collecting water quality data produced by different water related actors across France. Nutrient data was also collected directly with regional water agencies (<a href="https://www.eau-seine-normandie.fr/">https://www.eau-seine-normandie.fr/</a>, <a href="https://eau-grandsudouest.fr/">https://eau-grandsudouest.fr/</a>, <a href="https://www.eaurmc.fr/">https://www.eaurmc.fr/</a>, <a href="https://www.eau-artois-picardie.fr/">https://www.eau-artois-picardie.fr/</a>, <a href="https://www.eau-rhin-meuse.fr/">https://www.eau-rhin-meuse.fr/</a> and <a href="https://agence.eau-loire-bretagne.fr/home.html">https://agence.eau-loire-bretagne.fr/home.html</a>), and pre-processed using a database management system relying on PostgreSQL with PostGIS extension (Thieu & Silvestre, 2015). A three-pass strategy was used to curate raw carbon and nutrients data: 1. Removal of “obvious outliers”, 2. Detection of baseline change and correction if possible (or removal of data) 3. Removal of outliers using a quantile based approach by element and temporal series.</p> <p>Hydrological time series are interpolation trough hydrograph transfer (de Lavenne et al., 2023) of 1664 time series of discharge completed with GR4J model (Pelletier & Andréassian, 2020; Pelletier 2021).</p> <p>Land cover data was extracted from Corine Land Cover dataset for years 2000, 2006, 2012, and 2018 (EEA, 2020). Raw CLC typology contains 44 classes. Results of percent cover per year per class were computed for each catchment. An aggregated typology of 8 classes is also proposed.</p> <p>Climatological data was extracted from daily reconstruction at 5 arcmin for temperatures and 1 arcmin for precipitation over Europe (Thiemig et al., 2022). Mean by catchment for min&max daily temperature and precipitation were computed for each catchment for the 1990-2019 period.</p> <p><strong>Nuts-STeauRY dataset</strong></p> <p><strong>Carbon and nutrients time series</strong></p> <p>Time series of carbon and nutrients within the 1962-2019 period on 5470 stations: Dissolved Organic Carbon (DOC), Total Organic Carbon (TOC) Nitrates (NO3-), Nitrites (NO2-), Ammonia (NH4+), Soluble Reactive Phosphorus (SRP), Total Phosphorus (TP) and Dissolved Silica (DSi).</p> <p><code>|var | n_unique_station| n_total_meas| mean_duration_y| mean_frequency_y|</code></p> <p><code>|:---|----------------:|------------:|---------------:|----------------:|</code></p> <p><code>|DOC | 4 992| 658 147| 14.3| 9.0|</code></p> <p><code>|DSi | 3 299| 333 866| 12.9| 8.3|</code></p> <p><code>|NH4 | 5 318| 907 343| 19.3| 8.7|</code></p> <p><code>|NO2 | 5 264| 891 886| 19.2| 8.6|</code></p> <p><code>|NO3 | 5 465| 939 279| 19.0| 9.0|</code></p> <p><code>|SRP | 5 361| 910 107| 19.1| 8.7|</code></p> <p><code>|TOC | 935| 111 993| 13.6| 9.6|</code></p> <p><code>|TP | 5 199| 802 841| 17.1| 8.8|</code></p> <p>Note that some SRP and DSi measurements were declared as realized on raw water. A thorough analysis of time series show no evidence of difference on baselines. For more accuracy, it is advised to filter out those analyses using the “fraction” attribute of each measurement.</p> <p><strong>Discharge modelled daily time series</strong></p> <p>Modelled naturalized discharge through hydrograph transfer and interpolated measured discharges when available for the 1980-2019 period.</p> <p>A daily discharge was computed for 5128 catchments. For small catchments (< 1000 km<sup>2</sup>, n = 4530), hydrograph transfer was used, while for big catchments, a direct interpolation of measured/completed discharges was performed. The direct interpolation was only possible for 598 catchments > 1000 km<sup>2</sup>. The criteria retained for a direct interpolation is 0.8*area_discharge_station < area_quality < 1.2*area_discharge_station when discharge and quality stations were nested.</p> <p>Hydrological time series uncertainties varies a lot depending on: quality of data source, distance from pseudo-gauged outlets, land cover of the catchments, natural spatial and temporal variability of discharge, size of the catchment (de Lavenne et al., 2016). We advise a cautious use of those modelled discharges as uncertainties could not be computed.</p> <p><strong>Catchments, outlets and conditioned DEM</strong></p> <p>5470 catchments and outlets are delivered as geopackages (EPSG: 3035).</p> <p>The DEM, conditioned by CCM 2.1 is also delivered as a GeoTIFF (EPSG: 3035) as way to delimit new catchment for the area that are consistent with the dataset.</p> <p><strong>Catchments characteristics and climate</strong></p> <p>Refer to Data sources & processing and File descriptions.</p> <p><strong> </strong></p> <p><strong>File and attributes descriptions: </strong></p> <p>The key “sta_code” is present across all files. For time varying records, “date” can be a secondary key. </p> <p><strong>Description of CNPSi.csv data attributes</strong></p> <p>Each line is a couple measurement/parameter/station</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· var: Abbreviation of parameter name</p> <p>· fraction: "water_filtrated" or "water_raw"</p> <p>· date: date of sampling</p> <p>· hour: hour of sampling</p> <p>· value: analytical result (concentration)</p> <p>· provider: provider of the data</p> <p>· producer: producer of the data</p> <p>· from_db: "Naiades2022" (https://naiades.eaufrance.fr/france-entiere#/ dump from 2022) or "DoNuts" (Thieu, V., Silvestre, M., 2015. DoNuts: un système d’information sur les observations environnementales. Présentation Séminaire UMR Métis)</p> <p>· n_meas: number of observations for a given parameter / station</p> <p>· unit: unit of concentration</p> <p>· element: "C" "N" "P" or "Si"</p> <p>· year: year of observation</p> <p>· month: month of observation</p> <p>· day: day of observation</p> <p>· julian_day: julian day observation (1-366)</p> <p>· decade: decade of observation (one of "1961-1970", "1971-1980", "1981-1990", "1991-2000", "2001-2010", "2011-2020")</p> <p> </p> <p> </p> <p><strong>Description of CNPSi_stats.csv data attributes</strong></p> <p>Each line is a couple parameter / station</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· var: Abbreviation of parameter name</p> <p>· n_meas: number of observations for a given parameter / station</p> <p>· start_year: year of first observation for a given parameter / station</p> <p>· end_year: year of last observation for a given parameter / station</p> <p>· duration_y_tot: total duration of observation in years for a given parameter / station</p> <p>· duration_y_tot: duration of observation in years for a given parameter / station for years with at least 1 meas</p> <p>· mean_nmeas_per_y_tot: mean number of observations per year considering total duration</p> <p>· mean_nmeas_per_y_meas: mean number of observations per year considering years with measurements</p> <p>· is_fully_continuous: TRUE if at least one measurement per year for a given parameter / station</p> <p>· start_cont_seq: year in which starts the longest continuous sequence for a given parameter / station</p> <p>· end_cont_seq: year in which ends the longest continuous sequence for a given parameter / station</p> <p>· duration_y_cont_seq: duration in years for the longest continuous sequence for a given parameter / station</p> <p>· nmeas_cont_seq: number of measurements for the longest continuous sequence for a given parameter / station</p> <p>· mean_nmeas_per_y_cont_seq: mean number of observations per year for the longest continuous sequence for a given parameter / station</p> <p>· mean: mean value (concentration) for a given parameter / station</p> <p>· median: median value (concentration) for a given parameter / station</p> <p>· sd: standard deviation (concentration) for a given parameter / station</p> <p>· cv: coeficient of variation (concentration) for a given parameter / station</p> <p>· c05,c25,c50,c75,c95: centiles 5, 25, 50, 75 & 95 for a given parameter / station</p> <p><strong>Description of catchments.gpkg and outlets.gpkg data attributes</strong></p> <p>Each line is a catchment or an outlet (sampling point)</p> <p>File is a .gpkg (EPSG = 3035)</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· watercourse: Name of the water course (from spatial join on IGN BD Topo)</p> <p>· mun_name: Name of the municipality of the outlet (from spatial join on IGN BD Admin Express)</p> <p>· ccm_wso_id: Seaoutlet id from CCM v2.1 database</p> <p>· ccm_wso1_id: Elementary catchment id from CCM v2.1 database</p> <p>· ccm_strahler: Strahler order of the catchment from CCM v2.1 database</p> <p>· area_km2: Computed area in km2 of the catchment</p> <p><strong>Description of daily discharges data attributes</strong></p> <p>Each line corresponds to a daily modelled discharge at a quality station from 1980 to 2019</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· date: Date in format yyyy-mm-dd</p> <p>· flow_mm: Discharge expressed in mm.d-1</p> <p>· flow_m3s: Discharge expressed in m3.s-1</p> <p><strong>Description of climate data attributes</strong></p> <p>Each line in the pr_tmin_tmax_1990-2019_lt_mean.csv corresponds to a mean value within a catchment for the 1990-2019 period.</p> <p>· sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>· period: 1990-2019</p> <p>· source: EMO-1 (pr) & EMO-5 (tmin, tmax)</p> <p>· pr: mean yearly precipitation (mm)</p> <p>· tmin: mean daily minimal temperature (°C)</p> <p>· tmin: mean daily maximal temperature (°C)</p> <p><strong>Description of land cover data attributes</strong></p> <p>Each line in the clc_8class.csv and clc_44class.csv corresponds to Corine Land Cover (CLC) class for a year (1990, 2000, 2006, 2012, or 2018) and a catchment. Raw CLC typology describes 44 classes that were aggregated to 8 classes (see clc_44class_to_8class.csv).</p> <p>· clc_44class.csv</p> <p>o sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>o year: Year as stated in CLC product</p> <p>o clc_name: Description of land cover class in CLC product</p> <p>o clc_code: Code for land cover class in CLC product</p> <p>o percent_cover: Percent cover by CLC class in the catchment (0-100)</p> <p>· clc_8class.csv</p> <p>o sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>o year: Year as stated in CLC product</p> <p>o label_clc_8class: Description of land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p>o code_clc_8class: Code for land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p>o percent_cover: Percent cover by aggregated CLC class in the catchment (0-100)</p> <p>· clc_44class_to_8class.csv</p> <p>o code_clc: Code for land cover class in CLC product (44 classes)</p> <p>o code_clc_8class: Code for land cover class in aggregated CLC product (8classes)</p> <p>o label_clc_8class: Description of land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p> </p> <p> </p> <p><strong>Acknowledgement</strong></p> <p>This publication has been prepared using European Union's Copernicus Land Monitoring Service information; <a href="https://doi.org/10.2909/960998c1-1870-4e82-8051-6485205ebbac">https://doi.org/10.2909/960998c1-1870-4e82-8051-6485205ebbac</a></p> <p>The authors thank Vasken Andréassian for communicating the discharge data and discharge station data and Alban de Lavenne for its help in using the transfr package, both for INRAE UR HYCAR.</p> <p> </p> <p><strong>References</strong></p> <p>de Lavenne, A., Skøien, J. O., Cudennec, C., Curie, F., & Moatar, F. (2016). Transferring measured discharge time series: Large-scale comparison of Top-kriging to geomorphology-based inverse modeling: transferring measured discharge time series. Water Resources Research, 52(7), 5555–5576. https://doi.org/10.1002/2016WR018716</p> <p>de Lavenne, A., Loree, T., Squividant, H., & Cudennec, C. (2023). The transfR toolbox for transferring observed streamflow series to ungauged basins based on their hydrogeomorphology. Environmental Modelling & Software, 159, 105562. <a href="https://doi.org/10.1016/j.envsoft.2022.105562">https://doi.org/10.1016/j.envsoft.2022.105562</a></p> <p>EEA. (2020). Corine Land Cover édition 2018. CLC 2018. <a href="https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-corine">https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-corine</a></p> <p>Pelletier, A., & Andréassian, V. (2020). Hydrograph separation: An impartial parametrisation for an imperfect method. Hydrology and Earth System Sciences, 24(3), 1171–1187. <a href="https://doi.org/10.5194/hess-24-1171-2020">https://doi.org/10.5194/hess-24-1171-2020</a></p> <p>Pelletier, A. (2021). Complétion d'hydrogrammes avec le modèle GR4J - Note méthodologique. INRAE, UR HYCAR.</p> <p>Thiemig, V., Gomes, G. N., Skøien, J. O., Ziese, M., Rauthe-Schöch, A., Rustemeier, E., Rehfeldt, K., Walawender, J. P., Kolbe, C., Pichon, D., Schweim, C., and Salamon, P.: EMO-5: a high-resolution multi-variable gridded meteorological dataset for Europe, Earth Syst. Sci. Data, 14, 3249–3272, https://doi.org/10.5194/essd-14-3249-2022, 2022</p> <p>Thieu, V., Silvestre, M., 2015. DoNuts : un système d'information sur les observations environnementales. Présentation Séminaire UMR Métis</p> <p>Vogt, J., A. de Jager, E. Rimaviciute, W. Mehl, S. Foisneau, K. Bódis, J. Dusart, M.L. Paracchini, P. Haastrup, & C. Bamps. (2007). A pan-European river and catchment database. (European Commission. Joint Research Centre. Institute for Environment and Sustainability.). Publications Office. https://data.europa.eu/doi/10.2788/35907</p>
Synthetic Dataset of Citation Strings in 12 Styles
<p>This dataset was produced in the aim of testing different tools for citation string parsing, as part of the experiment reported in the paper:</p> <blockquote> <p>Iana Atanassova and Marc Bertin, 2024. "Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers", Bibliometric-enhanced Information Retrieval workshop (BIR), collocated with ECIR 2024, Glasgow, Scotland. </p> </blockquote> <h2><br>Data</h2> <p>The data that is provided here is organised as follows:</p> <ul> <li>the file <strong>citation-strings.zip</strong> contains raw citation strings that were generated for each of the 12 citation styles in txt format</li> <li>the file <strong>parsers-output.csv</strong> contains the output that was produced from the parsers: ChatGPT, Llama, and Neural ParsCit</li> </ul> <h2><br>To cite this work</h2> <p>To use this dataset and/or the results produced in the experiment, please cite the following article:</p> <blockquote> <p>@inproceedings{atanassova2024citparse,<br> title = {{Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers}}, <br> author = {Iana Atanassova and Marc Bertin},<br> year = {2024},<br> booktitle = {{International Workshop on Bibliometric-enhanced Information Retrieval (BIR 2024) co-located with the 46\textsuperscript{st} European Conference on Information Retrieval (ECIR 2024)}},<br> address = {Glasgow, Scotland}<br>}</p> </blockquote> <h3>Authors information</h3> <ul> <li>Iana Atanassova, ORCID https://orcid.org/0000-0003-3571-4006 URL https://iana-atanassova.github.io/</li> <li>Marc Bertin, ORCID https://orcid.org/0000-0003-1803-6952 URL https://elico-recherche.msh-lse.fr/membres/marc-bertin</li> </ul> <h3>Related github repository</h3> <p>https://github.com/iana-atanassova/citation-parsers-bir2024.git </p>
Power Balance Characteristics for Multirotor- and Fixed-Wing-Type UAV-BSs Equipped with RES and RISs
<h2><strong>Overview</strong></h2> <p>The following dataset presents the power balance characteristics for Unmanned Aerial Vehicle Base Stations (UAV-BSs) equipped with Renewable Energy Sources (RES) and Reconfigurable Intelligent Surfaces (RISs). The dataset has been prepared for two different types of UAVs, i.e., multirotor and fixed-wing ones.</p> <h2><strong>Scenario</strong></h2> <p>The considered scenario includes 2 UAV-BSs (each of a different type) equipped with a single RF transceiver and an RIS device and RES — a single photovoltaic panel (PV) and a single wind turbine (WT). The UAV-BSs are placed within the city of Poznan and hover (multirotor) or follow a circular route (fixed-wing) above a single mobile user with fixed traffic demand (100 Mbps downlink — DL, and 50 Mbps uplink — UL). The simulation runs have been performed for 4 dates (vernal equinox, summer solstice, autumn equinox, winter solstice), each one from a different season of the year. The aim of such an approach was to highlight the impact of the time of the day and the year on the energy gain obtained thanks to enabling RES generators as well as on the power consumption of the hardware of each UAV-BS type. The weather conditions assumed within the simulation are typical for the climate in Poland.</p> <h2><strong>Methodology</strong></h2> <p>The power-balance calculations (UAV-BSs' power consumption, renewable energy production) have been based on the mathematical formulas from the scientific literature and performed within the digital simulation runs by using dedicated software developed in Python programming language.</p> <h2><strong>Simulation setup</strong></h2> <p>The setup of the input parameters for used mathematical models (power consumption, energy generation) has been done in accordance with the values attached within the literature positions (cited within the publication included in the <em>Related works</em> section of the following dataset) and adjusted to the considered study. Furthermore, the data used to predict weather conditions are the real data (for the year 2022) collected by the weather stations placed in Poznan. A single simulation run has been performed (which takes into account 2 types of UAV-BS simultaneously and estimates their power balance for 4 seasons of the year), where the time step has been set to 1 hour of the day.</p> <h2><strong>Results</strong></h2> <p>The results of the aforementioned investigations have been included in the attached files (<em>_power_balance_multirotor.csv</em> & <em>_power_balance_fixed_wing.csv</em>). The first column denotes the hour of a particular day. Next, 4 multicolumns have been presented for the following variants — No RES enabled, only PV enabled, only WT enabled, and both types of RES generators enabled. In addition, each multicolumn consists of 4 columns, each of which represents a UAV-BS's hardware power balance (in W) for a different date (season of the year).</p> <h2><strong>Acknowledgment</strong></h2> <p>More details about the conducted study have been described within the attached paper (<em>Related works</em> section). The work (including the following dataset preparation) was realized within project no. 2021/43/B/ST7/01365 funded by the National Science Center in Poland.</p>
IPMWORKS Resource Toolbox - database extraction March 2024
<p><span>The IPMWORKS IPM Resource Toolbox (Toolbox) has been developed as an interactive, online repository of integrated pest management (IPM) resources. Populated with high priority resources for farmers and their advisors during the project, its structure enables additional resources added over time. The repository is a public interactive website, available to anyone looking to access, understand, and implement IPM. Built on an open-source content management system, the toolbox is designed to require minimal post-production site maintenance and support, while being easily expanded to integrate resources from future initiatives.<br>At the core of the Toolbox lies MongoDB, a powerful NoSQL database management system. The schema-less nature of MongoDB allows for flexible data modeling, crucial for accommodating the diverse array of materials within the IPMWORKS ecosystem. Additionally, the integration of GridFS, a feature of MongoDB, facilitates the storage and retrieval of large files like images, PDFs, and documents. This architectural choice ensures optimal performance and efficiency in handling a wide range of materials. <br>We here make available all content uploaded to the IPMWORKS Resource Toolbox up to 18 March 2024. Materials are available in two formats, first as MongoDB files which require users to open them as a Mongo database file, and second as JSON files. Note that the JSON format does not include access to any pdfs attached to Toolbox content, only the associated metadata. </span></p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.