Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

32,629

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

32,629 results for “Datasets”

Learn how ShareScore rates datasets ↗
zenodo52/100

TemStaPro Datasets

<p>This dataset contains protein sequences used to train, validate, and test binary classifiers that form TemStaPro program, which is applied&nbsp;for protein thermostability prediction with respect to nine&nbsp;temperature thresholds from&nbsp;40 to 80&nbsp;degrees Celsius using a step of five&nbsp;degrees.</p> <p>The data is given&nbsp;in files of FASTA format. Each protein sequence has a header made of three&nbsp;values separated by vertical bar&nbsp;symbols:&nbsp;organism's, to which the protein belongs, UniParc taxonomy identifier;&nbsp;UniProtKB/TrEMBL identifier of the protein sequence;&nbsp;organism's growth temperature taken from the dataset of growth temperatures of over 21 thousand organisms&nbsp;(Engqvist, 2018).</p> <p>TemStaPro-Major-30 set is composed of 12 files:</p> <ul> <li>one training</li> <li>one validation</li> <li>one imbalanced testing</li> <li>nine&nbsp;balanced samples&nbsp;of 2000 sequences from each of the balanced testing set</li> </ul> <p>TemStaPro-Minor-30 set is composed of cross-validation and testing files all balanced for 65 degrees Celsius temperature threshold.</p> <p>SupplementaryFileC2EPsPredictions.tsv file contains thermostability predictions using the default mode of TemStaPro program&nbsp;to check the thermostability of different C2EP groups.<br><br>The detailed description is given in the revised version of the corresponding paper (https://doi.org/10.1093/bioinformatics/btae157).</p> <p>If you use the data from this dataset, please cite both the paper and the DOI of the&nbsp;dataset.</p>

opencc-by-4.0Mar 2023View details →
zenodo52/100

Datasets of "Carbide coating on nickel to enhance the stability of supported metal nanoclusters" Nanoscale, 2022, 14, 3589-3598

<p>These are the datasets related to the publication &quot;Carbide coating on nickel to enhance the stability of supported metal nanoclusters&quot;, Nanoscale, 2022, 14, 3589-3598 (<a href="https://doi.org/10.1039/D1NR06485A">https://doi.org/10.1039/D1NR06485A</a>). They are saved as NeXus/HDF5 files according to the nxstm NeXus application definition (<a href="https://doi.org/10.5281/zenodo.5792930">https://doi.org/10.5281/zenodo.5792930</a>).</p>

opencc-by-4.0Aug 2022View details →
zenodo52/100

ColoPola: A dataset of colorectal cancer polarimetric images (Mueller matrix elements) for colorectal cancer detection

<p><strong>ColoPola</strong> dataset is <strong>Colo</strong>rectal cancer <strong>Pola</strong>rimetric images dataset</p> <p>The dataset consists of 572 slices (specimens) with 20,592 images, 284 slices of which were designated as cancer samples and 288 as normal samples.</p> <p>Each sample has 36 polarimetric images (i.e., HH, HV, HP, HM, HR, HL, VH, VV, VP, VM, VR, VL, PH, PV, PP, PM, PR, PL, MH, MV, MP, MM, MR, ML, RH, RV, RP, RM, RR, RL, LH, LV, LP, LM, LR, and LL).</p> <p>Each folder in the <strong>ColoPola</strong> dataset consists of 36 polarimetric images. Each image is 1280x1024 pixels in size and was created in the TIF file format (HH.tif, HV.tif, ..., LL.tif).&nbsp;</p>

opencc-zeroNov 2023View details →
zenodo52/100

BLASTNet Simulation Dataset

<p>Go to <a href="https://blastnet.github.io/">https://blastnet.github.io/</a>&nbsp;to access and download this&nbsp;reacting and non-reacting flow physics simulations.</p> <p><strong>Mission</strong></p> <p>BLASTNet 2.1 was developed to provide the researchers in &nbsp;reacting and non-reacting flow physics communities with high-fidelity simulation datasets in a convenient format for ML applications. With ~5 TB, 765 full-domain samples, and 36 configurations, BLASTNet can effectively address these gaps and aid in fostering open/fair ML development within reacting and non-reacting flow physics communities.</p> <p><strong>Application</strong></p> <p>This data is useful for fluid flows in a wide range of ML applications tied to automotive, propulsion, energy, and the environment. Specifically, scientific engineering tasks related to these domains may include turbulent closure modeling, spatio-temporal modeling, and inverse modeling.<br>&nbsp; &nbsp;&nbsp;</p>

opencc-by-4.0Jun 2023View details →
zenodo52/100

GlobalHighPM₂.₅: Global Daily Seamless 1 km Ground-Level PM₂.₅ Dataset over Land (2017–Present)

<p>GlobalHighPM<sub>2.5</sub> is part of a series of long-term, seamless, global, high-resolution, and high-quality datasets of air pollutants over land (i.e., GlobalHighAirPollutants, GHAP). It is generated from big data sources (e.g., ground-based measurements, satellite remote sensing products, atmospheric reanalysis, and model simulations) using artificial intelligence, taking into account the spatiotemporal heterogeneity of air pollution.</p> <p>This dataset contains input data, analysis codes, and generated dataset used for the following article. If you use the GlobalHighPM<sub>2.5</sub> dataset in your scientific research, please cite the following reference (Wei et al., NC, 2023):</p> <ul> <li> <p>Wei, J., Li, Z., Lyapustin, A., Wang, J., Dubovik, O., Schwartz, J., Sun, L., Li, C., Liu, S., and Zhu, T.&nbsp;<a href="https://weijing-rs.github.io/publications/Wei_et_al-NC-2023.pdf" target="_blank" rel="noopener">First close insight into global daily gapless 1 km PM<sub>2.5</sub>&nbsp;pollution, variability, and health impact</a>.&nbsp;<em>Nature Communications</em>, 2023, 14, 8349. https://doi.org/10.1038/s41467-023-43862-3</p> </li> </ul> <p><strong>Input Data</strong></p> <p>Relevant raw data for each figure (compiled into a single sheet within an Excel document) in the manuscript.</p> <p><strong>Code</strong></p> <p>Relevant Python scripts for replicating and ploting the analysis results in the manuscript, as well as codes for converting data formats.</p> <p><strong>Generated Dataset</strong></p> <p>Here is the first big data-derived seamless (spatial coverage = 100%) daily, monthly, and yearly 1 km (i.e., D1K, M1K, and Y1K) global ground-level PM<sub>2.5</sub> dataset over land from 2017 to the present. This dataset exhibits high quality, with cross-validation coefficients of determination (CV-R<sup>2</sup>) of 0.91, 0.97, and 0.98, and root-mean-square errors (RMSEs) of 9.20, 4.15, and 2.77 &micro;g m<sup>-3</sup>&nbsp;on the daily, monthly, and annual bases, respectively.</p> <p><strong>Due to data volume limitations,&nbsp;</strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; all (including <strong>daily</strong>) data for the year <strong>2022 </strong>is accessible at: <strong><a href="../records/10795661">GlobalHighPM2.5 (2022)</a></strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; all (including <strong>daily</strong>) data for the year <strong>2021 </strong>is accessible at: <strong><a href="../records/10398385">GlobalHighPM2.5 (2021)</a></strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; all (including <strong>daily</strong>) data for the year <strong>2020 </strong>is accessible at: <strong><a href="../records/10402639">GlobalHighPM2.5 (2020)</a></strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; all (including <strong>daily</strong>) data for the year <strong>2019 </strong>is accessible at: <strong><a href="../records/10402723">GlobalHighPM2.5 (2019)</a></strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; all (including <strong>daily</strong>) data for the year <strong>2018 </strong>is accessible at: <strong><a href="../records/10402824">GlobalHighPM2.5 (2018)</a></strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; all (including <strong>daily</strong>) data for the year <strong>2017 </strong>is accessible at: <strong><a href="../records/10403497">GlobalHighPM2.5 (2017)</a></strong></p> <p>&nbsp; &nbsp; &nbsp; &nbsp; continuously updated...</p> <p><strong>More GHAP datasets for different air pollutants are available at:&nbsp;<a href="https://weijing-rs.github.io/product.html">https://weijing-rs.github.io/product.html</a></strong></p>

opencc-by-4.0Apr 2022View details →
zenodo52/100

Dataset - Decrypting lysine deacetylase inhibitor action and protein modifications by dose-resolved proteomics

<h4><strong>Dataset Summary</strong></h4> <p>Lysine deacetylase inhibitors (KDACis) are approved for cutaneous T-cell lymphoma (CTCL), peripheral T-cell lymphoma (PTCL), and multiple myeloma. Despite the mechanism of action(s) (MoA) remains elusive, these inhibitors lead to increasing acetylation levels of histones and other proteins, altered gene expression and cell death. To characterize the MoA of these drugs in more detail, we systematically measured dose-dependent changes in protein expression, acetylation, and phosphorylation in response to 21 clinical and pre-clinical KDACis. MV4-11 cells were treated for 6 h with 1 vehicle control and 10 increasing doses of the respective drug (from 100 pM to 30 mM). Proteins were digested with trypsin, and the resulting 11 peptide preparations corresponding to one drug dose each were encoded by stable isotopes (tandem mass tags, TMT-11plex) and combined. Acetylated peptides were subsequently enriched by immunoprecipitation and phosphopeptides by immobilized metal affinity chromatography (IMAC). PTM-carrying and unmodified peptides were analyzed separately by liquid chromatography tandem mass spectrometry (LC-MS/MS) for peptide and protein identification and quantification. Additionally, Vorinostat and Panobinostat were also recorded as time-dependent experiments at their pEC50 concentration, respectively.&nbsp;</p> <h4><strong>Dataset structure</strong></h4> <p>Here, we provide all curve data processed with CurveCurator v0.4.0 (<a href="https://github.com/kusterlab/curve_curator">https://github.com/kusterlab/curve_curator</a>). Each drug is a zip folder containing acetylome, phosphoproteome, and fullproteome data. Next to each data set is the toml parameter file used to generate the curves.txt and dashboard.html files. Time-dependent data is indicated by "td" and dose-dependent data is indicated by "dd".</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo52/100

Dataset of "Fast carbon dioxide–epoxide cycloaddition catalyzed by metal and metal-free ionic liquids for designing non-isocyanate polyurethanes"

<p>The recycling of industrially produced greenhouse gases, such as CO2, into high-value-added chemicals is one of the most relevant strategies for reaching climate targets. A two-step strategy for designing non-isocyanate polyurethanes (NIPUs) from renewable carbon dioxide (CO2) using environmentally friendly conditions and catalysts is investigated. The first reaction step efficiently converts a mono-epoxidized monomer (phenyl glycidyl ether) into cyclic carbonates under mild reaction conditions and supercritical CO2, using imidazolium ionic liquids (ILs) as catalysts into cyclic carbonates. The DFT calculations suggested a comprehensive mechanistic pathway for the IL-catalyzed CO2-epoxy reaction showing a rate-determining step of the initial epoxide ring opening and the direct participation of IL-anions.</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Dataset for Accuracy of Grid-Connected Photovoltaic Power Plant: A Novel Approach Using Hybrid Variational Mode Decomposition and CNN-LSTM Model

<p>This research paper introduces a deep learning hybrid model employing Convolutional Neural Network Long Short-Term Memory (CNN-LSTM) for short-term photovoltaic (PV) solar energy forecasting.The proposed method integrates the Variational Mode Decomposition (VMD) algo-rithm with the CNN-LSTM model to predict PV power generation from a solar farm in Boussada, Algeria, from January 1, 2019, to December 31, 2020. The performance of the developed model is benchmarked against other deep learning models (VMD-CNN, VMD-LSTM, CNN-LSTM) across various time horizons (15, 30, and 60 minutes) to provide a comprehensive evaluation. Our findings exhibit greater performance of the developed model compared to other architectures, showcasing promising results in solar power forecasting. This research contributes to the main goal of enhancing EMS by providing accurate solar energy forecasts.</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

nuts-STeauRY dataset: hydrochemical and catchment characteristics dataset for large sample studies of Carbon, Nitrogen, Phosphorus and Silicon in french watercourses

<p><strong>nuts-STeauRY dataset: hydrochemical and catchment characteristics dataset for large sample studies of Carbon, Nitrogen, Phosphorus and Silicon in French watercourses</strong></p> <p>Antoine Casquin, Marie Silvestre, Vincent Thieu</p> <p>10.5281/zenodo.10830852</p> <p>v0.1, 18<sup>th</sup> March 2024</p> <p><strong>Brief overview of data: </strong></p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Carbon and nutrients data for 5470 continental French catchments</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Modelled discharge for 5128 of catchments out of 5470</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Geopackages with catchment delineations and outlets</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; DEM conditioned to delimit additional catchments</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Land-use and climatic data for 5470 continental French catchments</p> <p><strong>Citation of this work<br></strong></p> <p>A data paper is currently being submitted with details of methods and results. Once published, it will be the preferential source to cite. The data paper will be link to the new version of the dataset that will be updated on doi.org/10.5281/zenodo.10830852. If you use this dataset in your research or report, you must cite it.</p> <p><strong>Motivations</strong></p> <p>Data was collected and curated for the nuts-STeauRY project (<a href="http://nuts-steaury.cnrs.fr">http://nuts-steaury.cnrs.fr</a>), which deployed a national generic land to sea modelling chain.</p> <p>Data was primarily used (see related works):</p> <ol> <li>To calibrate concentrations of dissolved organic carbon and dissolve silica in headwaters</li> <li>To validate spatially and temporally the modelling chain (DOC, NO3-, NH4+, TP, SRP, DSi)</li> </ol> <p>Hydrochemical large sample datasets have numerous other uses: trends computations elucidate transfer mechanisms, machine learning, retrospective studies etc.</p> <p>The objective here is to provide a large sample curated dataset of carbon and nutrients concentrations along with modelled discharges, catchment characteristics and delimitations for the continental France. Such large sample dataset aims at easing the large sample studies over France and/or Europe. Although part of the data gathered here is obtainable via public sources, the catchments delineations, their characteristics and modelled hydrology were note not publicly available yet.&nbsp;Moreover, a unification of units and detection and removal of outliers was performed on carbon and nutrients data.</p> <p><strong>Data sources &amp; processing</strong></p> <p>Sampling points where snapped on the CCM database v2.1 (<a href="http://data.europa.eu/89h/fe1878e8-7541-4c66-8453-afdae7469221">http://data.europa.eu/89h/fe1878e8-7541-4c66-8453-afdae7469221</a>)(Vogt et al., 2007) and catchments were delineated using a 100m resolution Digital Elevation Model &nbsp;(DEM) conditioned by the hydrographic network and elementary catchments&rsquo; delineations of the CCM data v2.1. <strong>More than 6000 catchments were delineated and screened manually</strong> to check consistency: 5470 were retained<strong>.</strong></p> <p>Nutrient data was collected mainly through the Naiades portal (<a href="https://naiades.eaufrance.fr/">https://naiades.eaufrance.fr/</a>), a database collecting water quality data produced by different water related actors across France. Nutrient data was also collected directly with regional water agencies (<a href="https://www.eau-seine-normandie.fr/">https://www.eau-seine-normandie.fr/</a>, <a href="https://eau-grandsudouest.fr/">https://eau-grandsudouest.fr/</a>, <a href="https://www.eaurmc.fr/">https://www.eaurmc.fr/</a>, <a href="https://www.eau-artois-picardie.fr/">https://www.eau-artois-picardie.fr/</a>, <a href="https://www.eau-rhin-meuse.fr/">https://www.eau-rhin-meuse.fr/</a> and <a href="https://agence.eau-loire-bretagne.fr/home.html">https://agence.eau-loire-bretagne.fr/home.html</a>), and pre-processed using a database management system relying on PostgreSQL with PostGIS extension (Thieu &amp; Silvestre, 2015). A three-pass strategy was used to curate raw carbon and nutrients data: 1. Removal of &ldquo;obvious outliers&rdquo;, 2. Detection of baseline change and correction if possible (or removal of data) 3. Removal of outliers using a quantile based approach by element and temporal series.</p> <p>Hydrological time series are interpolation trough hydrograph transfer (de Lavenne et al., 2023) of 1664 time series of discharge completed with GR4J model (Pelletier &amp; Andr&eacute;assian, 2020; Pelletier 2021).</p> <p>Land cover data was extracted from Corine Land Cover dataset for years 2000, 2006, 2012, and 2018 (EEA, 2020). Raw CLC typology contains 44 classes. Results of percent cover per year per class were computed for each catchment. An aggregated typology of 8 classes is also proposed.</p> <p>Climatological data was extracted from daily reconstruction at 5 arcmin for temperatures and 1 arcmin for precipitation over Europe (Thiemig et al., 2022). Mean by catchment for min&amp;max daily temperature and precipitation were computed for each catchment for the 1990-2019 period.</p> <p><strong>Nuts-STeauRY dataset</strong></p> <p><strong>Carbon and nutrients time series</strong></p> <p>Time series of carbon and nutrients within the 1962-2019 period on 5470 stations: Dissolved Organic Carbon (DOC), Total Organic Carbon (TOC) Nitrates (NO3-), Nitrites (NO2-), Ammonia (NH4+), Soluble Reactive Phosphorus (SRP), Total Phosphorus (TP) and Dissolved Silica (DSi).</p> <p><code>|var | n_unique_station| n_total_meas| mean_duration_y| mean_frequency_y|</code></p> <p><code>|:---|----------------:|------------:|---------------:|----------------:|</code></p> <p><code>|DOC |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 4 992|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 658 147|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 14.3|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 9.0|</code></p> <p><code>|DSi |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 3 299|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 333 866|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 12.9|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;8.3|</code></p> <p><code>|NH4 |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 5 318|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 907 343|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 19.3|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 8.7|</code></p> <p><code>|NO2 |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 5 264|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 891 886|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 19.2|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 8.6|</code></p> <p><code>|NO3 |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 5 465|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 939 279|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 19.0|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 9.0|</code></p> <p><code>|SRP |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 5 361|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 910 107|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 19.1|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 8.7|</code></p> <p><code>|TOC |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 935|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 111 993|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 13.6|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 9.6|</code></p> <p><code>|TP&nbsp; |&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 5 199|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 802 841|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 17.1|&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; 8.8|</code></p> <p>Note that some SRP and DSi measurements were declared as realized on raw water. A thorough analysis of time series show no evidence of difference on baselines. For more accuracy, it is advised to filter out those analyses using the &ldquo;fraction&rdquo; attribute of each measurement.</p> <p><strong>Discharge modelled daily time series</strong></p> <p>Modelled naturalized discharge through hydrograph transfer and interpolated measured discharges when available for the 1980-2019 period.</p> <p>A daily discharge was computed for 5128 catchments. For small catchments (&lt; 1000 km<sup>2</sup>, n = 4530), hydrograph transfer was used, while for big catchments, a direct interpolation of measured/completed discharges was performed. The direct interpolation was only possible for 598 catchments &gt; 1000 km<sup>2</sup>. The criteria retained for a direct interpolation is 0.8*area_discharge_station &lt; area_quality &lt; 1.2*area_discharge_station when discharge and quality stations were nested.</p> <p>Hydrological time series uncertainties varies a lot depending on: quality of data source, distance from pseudo-gauged outlets, land cover of the catchments, natural spatial and temporal variability of discharge, size of the catchment (de Lavenne et al., 2016). We advise a cautious use of those modelled discharges as uncertainties could not be computed.</p> <p><strong>Catchments, outlets and conditioned DEM</strong></p> <p>5470 catchments and outlets are delivered as geopackages (EPSG: 3035).</p> <p>The DEM, conditioned by CCM 2.1 is also delivered as a GeoTIFF (EPSG: 3035) as way to delimit new catchment for the area that are consistent with the dataset.</p> <p><strong>Catchments characteristics and climate</strong></p> <p>Refer to Data sources &amp; processing and File descriptions.</p> <p><strong>&nbsp;</strong></p> <p><strong>File and attributes descriptions: </strong></p> <p>The key &ldquo;sta_code&rdquo; is present across all files. For time varying records, &ldquo;date&rdquo; can be a secondary key. &nbsp;</p> <p><strong>Description of CNPSi.csv data attributes</strong></p> <p>Each line is a couple measurement/parameter/station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; var: Abbreviation of parameter name</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; fraction: "water_filtrated" or "water_raw"</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; date:&nbsp; date of sampling</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; hour: hour of sampling</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; value: analytical result (concentration)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; provider: provider of the data</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; producer: producer of the data</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; from_db: "Naiades2022" (https://naiades.eaufrance.fr/france-entiere#/ dump from 2022) or "DoNuts" (Thieu, V., Silvestre, M., 2015. DoNuts: un syst&egrave;me d&rsquo;information sur les observations environnementales. Pr&eacute;sentation S&eacute;minaire UMR M&eacute;tis)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; n_meas: number of observations for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; unit: unit of concentration</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; element: "C" "N" "P" or "Si"</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; year: year of observation</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; month: month of observation</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; day: day of observation</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; julian_day: julian day observation (1-366)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; decade: decade of observation (one of "1961-1970", "1971-1980", "1981-1990", "1991-2000", "2001-2010", "2011-2020")</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Description of CNPSi_stats.csv data attributes</strong></p> <p>Each line is a couple parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; var: Abbreviation of parameter name</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; n_meas: number of observations for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; start_year: year of first observation for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; end_year: year of last observation for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; duration_y_tot: total duration of observation in years for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; duration_y_tot: duration of observation in years for a given parameter / station for years with at least 1 meas</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; mean_nmeas_per_y_tot: mean number of observations per year considering total duration</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; mean_nmeas_per_y_meas: mean number of observations per year considering years with measurements</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; is_fully_continuous: TRUE if at least one measurement per year for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; start_cont_seq: year in which starts the longest continuous sequence for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; end_cont_seq: year in which ends the longest continuous sequence for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; duration_y_cont_seq: duration in years for the longest continuous sequence for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; nmeas_cont_seq:&nbsp; number of measurements for the longest continuous sequence for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; mean_nmeas_per_y_cont_seq: mean number of observations per year for the longest continuous sequence for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; mean: mean value (concentration) for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; median: median value (concentration) for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sd: standard deviation (concentration) for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; cv: coeficient of variation (concentration) for a given parameter / station</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; c05,c25,c50,c75,c95: centiles 5, 25, 50, 75 &amp; 95 for a given parameter / station</p> <p><strong>Description of catchments.gpkg and outlets.gpkg data attributes</strong></p> <p>Each line is a catchment or an outlet (sampling point)</p> <p>File is a .gpkg (EPSG = 3035)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_name: Name of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; watercourse: Name of the water course (from spatial join on IGN BD Topo)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; mun_name: Name of the municipality of the outlet (from spatial join on IGN BD Admin Express)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ccm_wso_id: Seaoutlet id from CCM v2.1 database</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ccm_wso1_id: Elementary catchment id from CCM v2.1 database</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; ccm_strahler: Strahler order of the catchment from CCM v2.1 database</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; area_km2: Computed area in km2 of the catchment</p> <p><strong>Description of daily discharges data attributes</strong></p> <p>Each line corresponds to a daily modelled discharge at a quality station from 1980 to 2019</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; date: Date in format yyyy-mm-dd</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; flow_mm: Discharge expressed in mm.d-1</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; flow_m3s: Discharge expressed in m3.s-1</p> <p><strong>Description of climate data attributes</strong></p> <p>Each line in the pr_tmin_tmax_1990-2019_lt_mean.csv corresponds to a mean value within a catchment for the 1990-2019 period.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; period: 1990-2019</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; source: EMO-1 (pr) &amp; EMO-5 (tmin, tmax)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; pr: mean yearly precipitation (mm)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; tmin: mean daily minimal temperature (&deg;C)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; tmin: mean daily maximal temperature (&deg;C)</p> <p><strong>Description of land cover data attributes</strong></p> <p>Each line in the clc_8class.csv and clc_44class.csv corresponds to Corine Land Cover (CLC) class for a year (1990, 2000, 2006, 2012, or 2018) and a catchment. Raw CLC typology describes 44 classes that were aggregated to 8 classes (see clc_44class_to_8class.csv).</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; clc_44class.csv</p> <p>o&nbsp;&nbsp; sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>o&nbsp;&nbsp; year: Year as stated in CLC product</p> <p>o&nbsp;&nbsp; clc_name: Description of land cover class in CLC product</p> <p>o&nbsp;&nbsp; clc_code: Code for land cover class in CLC product</p> <p>o&nbsp;&nbsp; percent_cover: Percent cover by CLC class in the catchment (0-100)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; clc_8class.csv</p> <p>o&nbsp;&nbsp; sta_code: Code of the station in the Sandre referentiel (public french "dataverse" for water data)</p> <p>o&nbsp;&nbsp; year: Year as stated in CLC product</p> <p>o&nbsp;&nbsp; label_clc_8class: Description of land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p>o&nbsp;&nbsp; code_clc_8class: Code for land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p>o&nbsp;&nbsp; percent_cover: Percent cover by aggregated CLC class in the catchment (0-100)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; clc_44class_to_8class.csv</p> <p>o&nbsp;&nbsp; code_clc: Code for land cover class in CLC product (44 classes)</p> <p>o&nbsp;&nbsp; code_clc_8class: Code for land cover class in aggregated CLC product (8classes)</p> <p>o&nbsp;&nbsp; label_clc_8class: Description of land cover class in CLC product aggregated in 8 classes (see clc_44class_to_8class.csv)</p> <p>&nbsp;</p> <p>&nbsp;</p> <p><strong>Acknowledgement</strong></p> <p>This publication has been prepared using European Union's Copernicus Land Monitoring Service information; <a href="https://doi.org/10.2909/960998c1-1870-4e82-8051-6485205ebbac">https://doi.org/10.2909/960998c1-1870-4e82-8051-6485205ebbac</a></p> <p>The authors thank Vasken Andr&eacute;assian for communicating the discharge data and discharge station data and Alban de Lavenne for its help in using the transfr package, both for INRAE UR HYCAR.</p> <p>&nbsp;</p> <p><strong>References</strong></p> <p>de Lavenne, A., Sk&oslash;ien, J. O., Cudennec, C., Curie, F., &amp; Moatar, F. (2016). Transferring measured discharge time series: Large-scale comparison of Top-kriging to geomorphology-based inverse modeling: transferring measured discharge time series. Water Resources Research, 52(7), 5555&ndash;5576. https://doi.org/10.1002/2016WR018716</p> <p>de Lavenne, A., Loree, T., Squividant, H., &amp; Cudennec, C. (2023). The transfR toolbox for transferring observed streamflow series to ungauged basins based on their hydrogeomorphology. Environmental Modelling &amp; Software, 159, 105562. <a href="https://doi.org/10.1016/j.envsoft.2022.105562">https://doi.org/10.1016/j.envsoft.2022.105562</a></p> <p>EEA. (2020). Corine Land Cover &eacute;dition 2018. CLC 2018. <a href="https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-corine">https://www.eea.europa.eu/data-and-maps/data/copernicus-land-monitoring-service-corine</a></p> <p>Pelletier, A., &amp; Andr&eacute;assian, V. (2020). Hydrograph separation: An impartial parametrisation for an imperfect method. Hydrology and Earth System Sciences, 24(3), 1171&ndash;1187. <a href="https://doi.org/10.5194/hess-24-1171-2020">https://doi.org/10.5194/hess-24-1171-2020</a></p> <p>Pelletier, A. (2021). Compl&eacute;tion d'hydrogrammes avec le mod&egrave;le GR4J - Note m&eacute;thodologique. INRAE, UR HYCAR.</p> <p>Thiemig, V., Gomes, G. N., Sk&oslash;ien, J. O., Ziese, M., Rauthe-Sch&ouml;ch, A., Rustemeier, E., Rehfeldt, K., Walawender, J. P., Kolbe, C., Pichon, D., Schweim, C., and Salamon, P.: EMO-5: a high-resolution multi-variable gridded meteorological dataset for Europe, Earth Syst. Sci. Data, 14, 3249&ndash;3272, https://doi.org/10.5194/essd-14-3249-2022, 2022</p> <p>Thieu, V., Silvestre, M., 2015. DoNuts : un syst&egrave;me d'information sur les observations environnementales. Pr&eacute;sentation S&eacute;minaire UMR M&eacute;tis</p> <p>Vogt, J., A. de Jager, E. Rimaviciute, W. Mehl, S. Foisneau, K. B&oacute;dis, J. Dusart, M.L. Paracchini, P. Haastrup, &amp; C. Bamps. (2007). A pan-European river and catchment database. (European Commission. Joint Research Centre. Institute for Environment and Sustainability.). Publications Office. https://data.europa.eu/doi/10.2788/35907</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Synthetic Dataset of Citation Strings in 12 Styles

<p>This dataset was produced in the aim of testing different tools for citation string parsing, as part of the experiment reported in the paper:</p> <blockquote> <p>Iana Atanassova and Marc Bertin, 2024. "Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers", Bibliometric-enhanced Information Retrieval workshop (BIR), collocated with ECIR 2024, Glasgow, Scotland.&nbsp;</p> </blockquote> <h2><br>Data</h2> <p>The data that is provided here is organised as follows:</p> <ul> <li>the file <strong>citation-strings.zip</strong> contains raw citation strings that were generated for each of the 12 citation styles in txt format</li> <li>the file <strong>parsers-output.csv</strong> contains the output that was produced from the parsers: ChatGPT, Llama, and Neural ParsCit</li> </ul> <h2><br>To cite this work</h2> <p>To use this dataset and/or the results produced in the experiment, please cite the following article:</p> <blockquote> <p>@inproceedings{atanassova2024citparse,<br>&nbsp; &nbsp; title = {{Breaking Boundaries in Citation Parsing: A Comparative Study of Generative LLMs and Traditional Out-of-the-box Citation Parsers}},&nbsp;<br>&nbsp; &nbsp; author = {Iana Atanassova and Marc Bertin},<br>&nbsp; &nbsp; year = {2024},<br>&nbsp; &nbsp; booktitle = {{International Workshop on Bibliometric-enhanced Information Retrieval (BIR 2024) co-located with the 46\textsuperscript{st} European Conference on Information Retrieval (ECIR 2024)}},<br>&nbsp; &nbsp; address = {Glasgow, Scotland}<br>}</p> </blockquote> <h3>Authors information</h3> <ul> <li>Iana Atanassova, ORCID https://orcid.org/0000-0003-3571-4006 URL https://iana-atanassova.github.io/</li> <li>Marc Bertin, ORCID https://orcid.org/0000-0003-1803-6952 URL https://elico-recherche.msh-lse.fr/membres/marc-bertin</li> </ul> <h3>Related github repository</h3> <p>https://github.com/iana-atanassova/citation-parsers-bir2024.git&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Datasets for: Generalizing Monin-Obukhov Similarity Theory (1954) for Complex Atmospheric Turbulence, Stiperski and Calaf 2023, PRL

<p>Scaling variables for the generalized flux-variance scaling relations that include turbulence anisotropy. Dataset is a companion to the manuscript &nbsp;Stiperski, I., Calaf, M., 2023: Generalizing Monin-Obukhov similarity theory (1954) for complex atmospheric turbulence. Physical Review Letters, 130 (12), 124001,&nbsp; &nbsp;https://doi.org/10.1103/PhysRevLett.130.124001</p> <p>The dataset contains the turbulence statistics from 13 datasets:&nbsp; AHATS, Cabauw, CASES-99, METCRAX II campaign (NEAR&nbsp; and RIM towers), T-Rex campaign (Central tower - TRexC, West tower - TRexW) and i-Box measurement network (CCS-VF0 tower - i-Box0, CS-SF1 tower - i-Box1, CS-NF10 tower - i-Box10, CS-NF27 tower - i-Box27, CS-MT21 tower - i-BoxTop, im Hinteren Eis tower - imHint).</p> <p><br>Data are organized in csv files for each datasets and only contain high quality (for applied criteria see the Supplemental Material of the companion paper, https://journals.aps.org/prl/supplemental/10.1103/PhysRevLett.130.124001) data with 30 min averaging for unstable stratification and 1 min for stable stratification. Since the data were used for scaling, there is no reference to time, but the measurement height is provided as an additional variable.&nbsp;</p> <p>Meaning of variables:</p> <p>zeta - z/L where z is height above ground and L is the local Obukhov length</p> <p>SigmaU - $\overline{u'u'}/u_*$ scaled standard deviation of streamwise velocity, where $u_*$ is the local friction velocity</p> <p>SigmaU - $\overline{v'v'}/u_*$ scaled standard deviation of spanwise velocity</p> <p>SigmaU - $\overline{v'v'}/u_*$ scaled standard deviation of surface-normal velocity</p> <p>SigmaT - $\overline{T'T'}/T_*$ scaled standard deviation of sonic temperature, where $T_*$ is the local temperature scale</p> <p>SigmaEpsU - scaled dissipation rate of the streamwise velocity</p> <p>SigmaEpsW - scaled dissipation rate of the surface-normal velocity&nbsp;</p>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Shoreline series of the Doniños coastal system, NW Iberia (1945-2020): A Geospatial Dataset

<p>This repository stores shoreline data spanning from 1945 to 2020, derived from aerial photography and orthophotos, for the Doni&ntilde;os coastal system in NW Iberia. The shoreline indicator is defined as the boundary between vegetated dunes and bare beach sand. The methodology and dataset are detailed in the following publication:</p> <p><em><strong>Rita Gonz&aacute;lez-Villanueva, Marti&ntilde;o Pastoriza, Armand Hern&aacute;ndez, Rafael Carballeira, Alberto S&aacute;ez, Roberto Bao. "Primary drivers of dune cover and shoreline dynamics: A conceptual model based on the Iberian Atlantic coast." Geomorphology, Volume 423, 2023, 108556, ISSN 0169-555X. <a href="https://doi.org/10.1016/j.geomorph.2022.108556" target="_new">https://doi.org/10.1016/j.geomorph.2022.108556</a>.</strong></em></p> <p>The shoreline dataset is encapsulated in a single GEOJSON file: <code>SHORES_1945_2020.geojson</code>. This dataset encompasses the shorelines mapped from all available aerial data for the Doni&ntilde;os coastal system, on the Galician coast, NW Iberia, from 1945 to 2020. It comprises a total of 15 shorelines. The geospatial layer employs the ETRS89/UTM zone 29N coordinate system (EPSG: 25829).</p> <p><strong>SHORES_1945_2020.geojson</strong>: This layer presents the shorelines, where each feature is a MultiLinestring with the following attributes:</p> <ul> <li><code>objectid</code>: Identifier of the shoreline.</li> <li><code>date</code>: Date of the shoreline capture, in month/year format (mm/yyyy).</li> <li><code>SHAPE_Leng</code>: Length of the mapped shoreline, in meters.</li> <li><code>Source</code>: Source of the original image from which the shoreline was derived, including the Centro Nacional de Informaci&oacute;n Geogr&aacute;fica (CNIG) and the Centro Cartogr&aacute;fico y Fotogr&aacute;fico del Ej&eacute;rcito del Aire (CECAF).</li> <li><code>Type</code>: 'O' signifies an orthophotograph, and 'AP' signifies an aerial photograph.</li> <li><code>GEO_Error</code>: Georeferencing error for each manually georeferenced photograph.</li> <li><code>geometry</code>: The type of geometry used in the file, specified as MultiLineString.</li> <li><code>coordinates</code>: UTM coordinates for each node in the multiline.</li> </ul>

opencc-by-4.0Mar 2024View details →
zenodo52/100

Dataset for "The influence of the amount of recycled material on the microstructure and properties of the second generation of single-domain YBCO bulks"

<p>The development of a recycling process for various REBCO materials is crucial considering both environmental sustainability and economic efficiency, particularly in light of the upcoming large-scale applications. In this paper, a novel general recycling process based on chemical dissolution was employed to grow REBCO bulks; recycled material obtained by recycling defective YBCO single-domain bulks was added (15 wt. %, 30 wt. % and 45 wt. %) to raw materials to prepare recycled YBCO precursor powder. Subsequently, recycled single-domain YBCO bulks were produced using Top-Seeded Melt Growth. The waste recycling related to of single-domain bulks growth was chosen, as it represents the most challenging form of waste in the context of REBCO superconductor production. The properties and microstructure of recycled bulks were further analyzed to determine the influence of the amount of recycled material used and compared to commercially produced bulks. Single-domain YBCO bulks were grown successfully from the recycled precursor powder. Furthermore, it was found that their properties could be tuned by varying the amount of the added recycled powder, allowing the use of vast amounts of REBCO waste for the preparation of bulks, when achieving the best possible properties is not essential for a given application. Given that the underlying recycling process is designed to work for all REBCO systems and any form of waste, it has significant implications for the sustainability and cost-effectiveness of REBCO superconductor production.&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Dataset to Schiedung et al. (2024): Millennial-aged pyrogenic carbon in high-latitude mineral soils

<p>Dataset to Schiedung et al. (2024, Communications Earth &amp; Environment): Pyrogenic Carbon is Aged at Millennial Scale in High-Latitude Mineral Soils</p> <p>DOI: <a href="https://doi.org/10.1038/s43247-024-01343-5">10.1038/s43247-024-01343-5</a></p> <p>This repository includes the following files:&nbsp;</p> <p><strong><em>dd_all.csv</em> </strong>- Includes all data for the individual samples that are presented in the manuscript.</p> <p><strong><em>Var_names_dd_all.csv</em> </strong>- Describes all variables in <em>dd_all</em> with corresponding unit&nbsp;</p> <p><strong><em>dd_site_average.csv</em></strong> - Includes all data that has been determined on composite samples for each site or the average of all samples per site&nbsp;</p> <p><strong><em>Var_names_dd_site_average.csv</em></strong>&nbsp; -&nbsp; Describes all variables in <em>dd_site_average.csv</em> with corresponding unit</p> <p>All .csv use "," as separator.&nbsp;</p> <p>This data set is also connected to Schiedung et al. (2022, Catena <a href="https://doi.org/10.1016/j.catena.2022.106194">&nbsp;https://doi.org/10.1016/j.catena.2022.106194</a> ) and the corresponding repository: <a href="../records/10609291">https://zenodo.org/records/10609291</a></p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Labelled acoustic dataset of roding Eurasian Woodcock (Scolopax rusticola)

<p>This dataset contains manually labelled audio data of roding Eurasian Woodcock&nbsp; (<em>Scolopax rusticola</em>).&nbsp;</p> <h2><strong>Description</strong></h2> <p>Bioacoustic surveys of roding Eurasian Woodcock were conducted in Baden-W&uuml;rttemberg, Germany in May and June in 2020 and 2021. The audio data of this collection was used for the evaluation of BirdNET as a means for the automated analysis of large quantities of audio data. The original dataset consisted of 12.236 minutes of recording, which were reviewed manually. Each call element of a male roding Woodcock (i.e. croak, whistle, chasing male) was annotated. Individual call elements were subsequently clustered into so called roding events, which are ecologically more meaningful. BirdNET was then tested against this manually labelled dataset.</p> <p>The dataset uploaded to zenodo contains:</p> <ul> <li>audio data of 2545 woodcock call element selections with a duration of 145 minutes</li> <li>audio data of 782 aggregated woodcock roding events with a duration of 115 minutes&nbsp;</li> <li>selection tables for call elements and roding events</li> <li>associated metadata</li> </ul> <p>Audio information in between roding events (i.e. non woodcock audio) ist not included due to data privacy reasons (see below).&nbsp;</p> <h3>Selections</h3> <p>Woodcock call elements were manually selected/annotated in Raven Pro with bounding boxes. For this dataset, all selections with a duration of less than 3 seconds were extended symmetrically until 3 seconds were reached. This may result in overlapping selections in the case of croaks that are directly followed by a whistle. Signals at the beginning or end of these selections may thus be included twice.</p> <h3>Roding events</h3> <p>A roding event was defined as a continuous series of Woodcock call elements with a maximum gap of six seconds between consecutive elements. Each event can be interpreted as a roding bird that passes by the recording location, similar to a typical woodcock roding survey conducted by a human observer. Roding events were not created with the extended 3 seconds clips described above, but with the original bounding box selections drawn in Raven Pro.</p> <h3>Audio files</h3> <ul> <li>selections.zip: each wav-file contains a single selections. Filenames correspond to the column selec in the table&nbsp;<em>selections.csv</em></li> <li>events.zip: each wav-file contains a single roding event, typically consisting of multiple call elements (croaks and/or whistles). In the case of faint signals of distant birds, roding events may consist of a single call element only. Filenames correspond to the column <em>event.id</em> in the table<em> events.csv</em>.</li> </ul> <h2><strong>Data collection</strong></h2> <p>All wav-files in this dataset originate from audio files that were recorded with autonomous recording units of the type AudioMoth. ARUs were housed in&nbsp; custom made waterproof casings (See details and files for 3D-printing: https://www.thingiverse.com/thing:6428228). ARUs were programmed to record continuously for 2 hours during dusk and were placed at edges of forest clearings. The devices were mounted to tree trunks at a height of approximately 1.5m above ground.&nbsp;</p> <h2><strong>Metadata files</strong></h2> <table> <tbody> <tr> <td><strong>filename</strong></td> <td><strong>content</strong></td> </tr> <tr> <td>sites.csv</td> <td> <p>contains locations of the recording sites. Since exact recording locations can not be made public, only recording sites (= cells of the 1km&sup2; UTM-grid) are provided. CRS: EPSG - 25832, ETRS89 / UTM 32N&nbsp;</p> <p>Data source of the underlying ETRS89 UTM 32N grid: https://gdz.bkg.bund.de/index.php/default/digitale-geodaten/nicht-administrative-gebietseinheiten/geographische-gitter-fur-deutschland-in-utm-projektion-geogitter-national.html</p> <p><strong>columns</strong></p> <p>site.id = unique id of recording sites,</p> <p>cellcode = official cellcode of the 1km&sup2;-UTM-grid</p> <p>elevation = mean elevation a.s.l.</p> <p>x.centroid = x-coordinate of centroid (EPSG: 25832)</p> <p>y.centroid = y-coordinate of centroid (EPSG: 25832)</p> <p>wkt.geometry = polygon geometry of the grid cell</p> </td> </tr> <tr> <td>arus.csv</td> <td> <p>metadata of the recording hardware</p> <p>&nbsp;</p> <p><strong>columns</strong></p> <p>aru.id = unique id of recording device</p> <p>type = recorder type</p> <p>manufacturer = manufacturer of recording hardware</p> <p>hardware.version = hardware version of the recording device</p> <p>acquisition.date = date the device was purchased (for reasons of microphone degradation)</p> </td> </tr> <tr> <td>deploys.csv</td> <td> <p>information on recorder deployment, includes aru settings, location, recording times&nbsp;</p> <p>&nbsp;</p> <p><strong>columns</strong></p> <p>deploy.id = unique id of recorder deployment</p> <p>aru.id = unique id of deployed aru</p> <p>start.date = date the aru was deployed in the field (YYYY-MM-DD)</p> <p>end.date = date the aru was collected (YYYY-MM-DD)</p> <p>firmware = firmware version used in this deployment</p> <p>rec.periods = number of daily recording periods (corresponds to start.rec1, start.rec2 ...)</p> <p>sample.rate = sample rate in kHz</p> <p>gain = gain setting</p> <p>sleep.duration = duration off stand-by phases in seconds, when set on a sleep/record-cycle</p> <p>rec.duration = duration of each recording in seconds, when set on a sleep/record-cycle</p> <p>start.rec1 = start of first recording period (UTC, hh:mm:ss)</p> <p>end.rec1 = end of first recording period (UTC, hh:mm:ss)</p> <p>start.rec2 = start of secondrecording period (UTC, hh:mm:ss)</p> <p>end.rec2 = end of second recording period (UTC, hh:mm:ss)</p> <p>site.id = unique id of recording site</p> </td> </tr> <tr> <td>recordings.csv</td> <td> <p>metadata of the audio files from which the roding events originate</p> <p>&nbsp;</p> <p>&nbsp;<strong>columns</strong></p> <p>recording.id = unique id of the recording</p> <p>deploy.id = unique id of aru deployment, during which the recording was made</p> <p>date = date on which the recording was made (YYYY-MM-DD)</p> <p>time = time of day at which the recording started (UTC, hh:mm:ss)</p> <p>duration = duration in seconds</p> <p>sampler.rate = sample rate in kHz</p> <p>channels = number of channels</p> <p>bits = bit depth</p> <p>samples = number of audio samples</p> <p>gain = gain setting of the aru</p> <p>voltage = battery voltage of the aru during recording</p> <p>temperature = ambient temperature during recording&nbsp;</p> <p>reviewer = anonymous id of staff who reviewed the file and annotated calls</p> <p>&nbsp;</p> </td> </tr> <tr> <td>selections.csv</td> <td> <p>manually labelled woodcock call elements (i.e. croaks, whistles, chases). Short selections were extended to 3 seconds by symmetrically adding time before and after the original selection. In the format of raven pro selection tables.</p> <p>&nbsp;</p> <p>&nbsp;<strong>columns</strong></p> <p>selec = unique id of the selection. Corresponds to the filename of the wav-files in the archive <em>selections.zip</em></p> <p><em>deploy.id = unique id of the aru deployment during which the roding event was recorded</em></p> <p>channel = audio channel</p> <p>start = start of the event in seconds from the start of the recording</p> <p>end = end of the event in seconds from the start of the recording</p> <p>bottom.freq = bottom frequency of the annotation bounding box</p> <p>top.frequency = top frequency of the annotation bounding box</p> <p>species.code = species code as used by BirdNET</p> <p>common.name = English common name as used by BirdNET</p> <p>annotation = contains annotations of call elements that are pooled in the roding event. Thus typcally equal to the number of annotated call element&nbsp;</p> <p>recording.id = id of the recording this roding eventoriginates from</p> </td> </tr> <tr> <td>events.csv</td> <td> <p>aggregated roding events consisting of contiuous sequences of manually labelled call elements. In the format of raven pro selection tables</p> <p>&nbsp;</p> <p>&nbsp;<strong>columns</strong></p> <p>event.id = unique id of roding event. Corresponds to the filename of the wav-files in the archive <em>events.zip&nbsp;</em></p> <p>channel = audio channel</p> <p>start = start of the event in seconds from the start of the recording</p> <p>end = end of the event in seconds from the start of the recording</p> <p>bottom.freq = bottom frequency of the annotation bounding box</p> <p>top.frequency = top frequency of the annotation bounding box</p> <p>species.code = species code as used by BirdNET</p> <p>common.name = English common name as used by BirdNET</p> <p>annotation = contains annotations of call elements that are pooled in the roding event. Thus typcally equal to the number of annotated call element&nbsp;</p> <p>recording.id = id of the recording this roding eventoriginates from</p> <p>deploy.id = unique id of the aru deployment during which the roding event was recorded</p> </td> </tr> <tr> <td>removed_audio_files.txt</td> <td>selection ids and event ids of audio files that were deleted because they included voices. Their metadata is still included in the files described above</td> </tr> </tbody> </table> <p>&nbsp;</p> <h2><strong>Data privacy</strong></h2> <p>Selections and roding events were checked for human voices and audio information was removed, in case it contained any. Audio segments that did not contain woodcock calls were not completely checked for human voices&nbsp; and can thus not be made available.</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Dataset of "Perovskite QDs embedded in polymer as a wavelength-shifting layer for UV-sensitized silicon sensors"

<p>Detection of UV radiation is becoming increasingly important for many applications. Here we present novel UV sensor construction on the basis of standard Si detector modification. Wavelength shifting mechanism is achieved by the luminescence effect of perovskite quantum dots embedded in polymer layers. We comprehensively characterize these composite materials, various sensor modification routes and the &nbsp;optical properties of UV-enhanced visible sensors. Modified S1227-16 BG silicon photodiodes and S13360-1375 CS MPPC photodetectors with the enhanced UV response are successively manufactured.micrographs; cross sections of model simulation or prediction (MSP).</p>

opencc-by-4.0Apr 2024View details →
zenodo52/100

Multilingual news article similarity dataset

<p>This dataset contains the extended version of the authors' earlier work:&nbsp;<a href="../records/6507872">https://zenodo.org/records/6507872,</a> where pairs of news articles drawn from the first half of 2020 are annotated for seven aspects of similarity in the original version as well as an additional FRAME aspect:</p> <ul> <li><strong>GEO</strong>:&nbsp;How similar is the geographic focus (places, cities, countries, etc.) of the two articles?</li> <li><strong>ENT:</strong>&nbsp;How similar are the named entities (e.g., people, companies, organizations, products, named living beings), excluding previously considered locations appearing in the two articles?</li> <li><strong>TIME</strong>&nbsp;Are the two articles relevant to similar time periods or describing similar time periods?</li> <li><strong>NAR</strong>&nbsp;How similar are the narrative schemas presented in the two articles?</li> <li><strong>OVERALL</strong>&nbsp;Overall, are the two articles covering the same substantive news story? (excluding style, framing, and tone)</li> <li><strong>STYLE</strong>&nbsp;Do the articles have similar writing styles?</li> <li><strong>TONE</strong> Do the articles have similar tones?</li> <li><strong>FRAME</strong> Do the articles have similar framing and express similar opinions?</li> </ul>

opencc-by-4.0Jan 2024View details →
zenodo52/100

A collection of datasets for software vulnerability detection

<p>This is a collection of datasets that are used for AI-based software vulnerability detection. All the datasets are in the .csv format and each row represents a sample. Each dataset includes a set of functions written in C and the target of each function is either 0 (non-vulnerable) or 1 (vulnerable).</p> <ol> <li><strong>data_C_Lin2017_test.csv:</strong> <ul> <li>Reference paper: <a href="https://dl.acm.org/doi/10.1145/3133956.3138840">Vulnerability Discovery with Function Representation Learning from Unlabeled Projects</a>, 2017.</li> <li>Data source on GitHub: <a href="https://github.com/DanielLin1986/function_representation_learning">https://github.com/DanielLin1986/function_representation_learning</a></li> <li>This dataset includes 44 vulnerable and 577 non-vulnerable functions from the LibPNG project.</li> </ul> </li> <li><strong>data_C_LineVul_test.csv:</strong> <ul> <li>Reference paper: <a href="https://ieeexplore.ieee.org/document/9796256">LineVul: A Transformer-based Line-Level Vulnerability Prediction</a>, 2022.</li> <li>Data source on Hugging Face: <a href="https://huggingface.co/datasets/Partha117/LineVul_Test_Dataset">https://huggingface.co/datasets/Partha117/LineVul_Test_Dataset</a></li> <li>This dataset includes 1055 vulnerable and 17809 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_PrimeVul_test.csv:</strong> <ul> <li>Reference paper: <a href="https://arxiv.org/abs/2403.18624">Vulnerability Detection with Code Language</a><br><a href="https://arxiv.org/abs/2403.18624">Models: How Far Are We?</a> 2024.</li> <li>Data source on GitHub: <a href="https://github.com/DLVulDet/PrimeVul">https://github.com/DLVulDet/PrimeVul</a></li> <li>From the data source, the primevul_test.jsonl was used to created this dataset.</li> <li>This dataset includes&nbsp;695 vulnerable and 25213 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_Choi2017_test.csv:</strong> <ul> <li>Reference paper: <a href="https://www.ijcai.org/proceedings/2017/0214.pdf">End-to-End Prediction of Buffer Overruns from Raw Source Code</a><br><a href="https://www.ijcai.org/proceedings/2017/0214.pdf">via Neural Memory Networks</a>, 2017.</li> <li>Data source on GitHub: <a href="https://github.com/mjc92/buffer_overrun_memory_networks">https://github.com/mjc92/buffer_overrun_memory_networks</a></li> <li>From GitHub, all the data in trainnig_100.txt, test_1_100.txt, test_2_100.txt,test_3_100.txt,test_4_100.txt, and corresponding _labels.txt files are combined to create this dataset.</li> <li>This dataset includes 7054 vulnerable and 6946 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_Devign_test.csv:</strong> <ul> <li>Reference paper: <a href="https://proceedings.neurips.cc/paper_files/paper/2019/file/49265d2447bc3bbfe9e76306ce40a31f-Paper.pdf">Devign: Effective Vulnerability Identification by Learning Comprehensive Program Semantics via Graph Neural Networks</a>, 2019</li> <li>Data source on Hugging Face: <a href="https://huggingface.co/datasets/claudios/code_x_glue_devign">https://huggingface.co/datasets/claudios/code_x_glue_devign</a></li> <li>From Hugging Face, all the data in train, validation, and test are combined to create this dataset.</li> <li>This dataset includes&nbsp;12460 vulnerable and 14858 non-vulnerable functions.</li> </ul> </li> <li><strong>data_C_Ours_{train,test}.csv:</strong> <ul> <li>This dataset is manually collected from projects on GitHub that have registered CVEs into NVD from 2002 to 2023. The 6,766 non-vulnerable code functions are extracted from the <a href="https://dl.acm.org/doi/10.1145/3607199.3607242">DiverseVul dataset</a> to increase the code diversity.&nbsp;</li> <li>This training set includes 5413 vulnerable and 5413 non-vulnerable functions.</li> <li>The test set includes 1353 vulnerable and 1353 non-vulnerable functions.</li> </ul> </li> </ol>

openmit-licenseApr 2024View details →
zenodo52/100

Soundscape Attributes Translation Project (SATP) Dataset

<p>The data and audio included here were collected for the Soundscape Attributes Translation Project (SATP). First introduced in Aletta et. al. (<a href="https://biblio.ugent.be/publication/8695720/file/8695735.pdf">2020</a>), the SATP is an attempt to provide validated translations of soundscape attributes in languages other than English. The recordings were&nbsp;used for headphones - based listening experiments.</p> <p>The data are provided to accompany publications resulting from this project and to provide a unique dataset of 1000s of perceptual responses to a standardised set of urban soundscape recordings. This dataset is the result of efforts from hundreds of researchers, students, assistants, PIs, and participants from institutions around the world. We have made an attempt to list every contributor to this Zenodo repo; if you feel you should be included, please get in touch.</p> <p><strong>Citation</strong>: If you use the SATP dataset or part of it, please cite our paper describing the data collection and this dataset itself.</p> <p><strong>Overview</strong>:&nbsp;The SATP dataset consists of 27 30-sec binaural audio recordings made in urban public spaces in London and one 60 sec stereo calibration signal.</p> <p>The recordings were made at locations as reported in Table 1 of the README.md (<strong>Recording locations</strong>),&nbsp;at various times of day by an operator wearing a binaural&nbsp;kit consisting of&nbsp;BHS II microphones and a SQobold (HEAD acoustics) device. Recordings were&nbsp;then exported to WAV via the ArtemiS SUITE software, using the original dynamic range from HDF. The listening experiment and the calibration procedure were intended for a headphone playback system (Sennheiser HD650 or similar open-back headphones recommended).&nbsp;</p> <p>The recordings were selected from an initial set of 80 recordings through a pilot study to ensure the test set had an even coverage of the soundscape circumplex space. These recordings were sent to the partner institutions (see Table 2 of the README.md) and assessed by approximately 30 participants in the institution's target language. The questionnaire used in each assessment is a translation of Method A Questionnaire, ISO 12913-2:2018. Each institution carried out their own lab experiment to collect data, then submitted their data to the team at UCL to compile into a single dataset. Some institutions included additional questions or translation options; the combined dataset (`SATP Dataset v1.x.xlsx`) includes only the base set of questions, the extended set of questions from each institution is included in the `Institution Datasets` folder.</p> <p>In all, SATP Dataset v1.4 contains 19,089 samples, including 707 participants, for 27 recordings, in 18 languages with contributions from 29 institutions.</p> <p><strong>Descriptions of the recordings, including GPS coordinates and sound sources, can be found in the README.md file.</strong></p> <p><strong>Format</strong>:&nbsp;The audio recordings are provided as 24 bit, 48 kHz, stereo WAV files. The combined dataset and Institutional datasets are provided as long tidy data tables in .xlsx files.</p> <p><strong>Calibration:&nbsp;</strong>The recommended calibration&nbsp;approach was based on the open-circuit voltage (OCV) procedure which was considered most&nbsp;accessible but other&nbsp;calibration procedures are also possible (Lam et. al. (<a href="https://arxiv.org/abs/2207.12899">2022</a>)). The provided calibration file is a computer generated sine wave at 1kHz, matching&nbsp;a&nbsp;sine wave recorded using the exact same setup at&nbsp;SPL of 94 dB. In case of the calibration signal&nbsp;playback level set to match&nbsp;SPL of 94 dB at the eardrum, all the 27 samples should be reproduced at realistic loudness.&nbsp;More details on OCV calibration procedure and other options&nbsp;you can find in&nbsp;Lam et. al. (<a href="https://arxiv.org/abs/2207.12899">2022</a>) and the attached documentation. PLEASE DO NOT EXPOSE YOURSELF NOR THE PARTICIPANTS TO THE CALIBRATION SIGNAL SET AT THE REALISTIC LEVEL AS IT&nbsp;CAN CAUSE HARM.</p> <p><strong>License and reuse</strong>:&nbsp;All SATP recordings are provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) License and are free to use. We encourage other researchers to replicate the SATP protocol and contribute new languages to the dataset. We also encourage the use of these recordings and the perceptual data for further soundscape research purposes. Please provide the proper attribution and get in touch with the authors if you would like to contribute a new translation or for any other collaborations.</p>

opencc-by-4.0Sep 2022View details →
zenodo52/100

Open AI Literature 2010-2020 Dataset

<p>The OAIL_10-20 dataset is comprised of OpenAlex records which reproduce the majority of the Web of Science (WoS) records analysed in the course of writing the paper Patterns in the Growth and Thematic Evolution of Artificial Intelligence Research: A Study Using Bradford Distribution of Productivity and Path Analysis, Gupta et al.</p> <p>This paper aims to utilise the Bradford distribution to provide a focused analysis of the thematic evolution of research patterns and growth, and applies this analysis to a corpus of AI papers published over the 10 years between 2010 and 2020.&nbsp;</p> <p>We provide this dataset to allow for researchers to reproduce the findings using open science.</p>

opencc-by-4.0Apr 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record