Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
635
datasets available to search
ShareScore release 0.9.0
Dataset results
635 results for “attribution”
Soundscape Attributes Translation Project (SATP) Dataset
<p>The data and audio included here were collected for the Soundscape Attributes Translation Project (SATP). First introduced in Aletta et. al. (<a href="https://biblio.ugent.be/publication/8695720/file/8695735.pdf">2020</a>), the SATP is an attempt to provide validated translations of soundscape attributes in languages other than English. The recordings were used for headphones - based listening experiments.</p> <p>The data are provided to accompany publications resulting from this project and to provide a unique dataset of 1000s of perceptual responses to a standardised set of urban soundscape recordings. This dataset is the result of efforts from hundreds of researchers, students, assistants, PIs, and participants from institutions around the world. We have made an attempt to list every contributor to this Zenodo repo; if you feel you should be included, please get in touch.</p> <p><strong>Citation</strong>: If you use the SATP dataset or part of it, please cite our paper describing the data collection and this dataset itself.</p> <p><strong>Overview</strong>: The SATP dataset consists of 27 30-sec binaural audio recordings made in urban public spaces in London and one 60 sec stereo calibration signal.</p> <p>The recordings were made at locations as reported in Table 1 of the README.md (<strong>Recording locations</strong>), at various times of day by an operator wearing a binaural kit consisting of BHS II microphones and a SQobold (HEAD acoustics) device. Recordings were then exported to WAV via the ArtemiS SUITE software, using the original dynamic range from HDF. The listening experiment and the calibration procedure were intended for a headphone playback system (Sennheiser HD650 or similar open-back headphones recommended). </p> <p>The recordings were selected from an initial set of 80 recordings through a pilot study to ensure the test set had an even coverage of the soundscape circumplex space. These recordings were sent to the partner institutions (see Table 2 of the README.md) and assessed by approximately 30 participants in the institution's target language. The questionnaire used in each assessment is a translation of Method A Questionnaire, ISO 12913-2:2018. Each institution carried out their own lab experiment to collect data, then submitted their data to the team at UCL to compile into a single dataset. Some institutions included additional questions or translation options; the combined dataset (`SATP Dataset v1.x.xlsx`) includes only the base set of questions, the extended set of questions from each institution is included in the `Institution Datasets` folder.</p> <p>In all, SATP Dataset v1.4 contains 19,089 samples, including 707 participants, for 27 recordings, in 18 languages with contributions from 29 institutions.</p> <p><strong>Descriptions of the recordings, including GPS coordinates and sound sources, can be found in the README.md file.</strong></p> <p><strong>Format</strong>: The audio recordings are provided as 24 bit, 48 kHz, stereo WAV files. The combined dataset and Institutional datasets are provided as long tidy data tables in .xlsx files.</p> <p><strong>Calibration: </strong>The recommended calibration approach was based on the open-circuit voltage (OCV) procedure which was considered most accessible but other calibration procedures are also possible (Lam et. al. (<a href="https://arxiv.org/abs/2207.12899">2022</a>)). The provided calibration file is a computer generated sine wave at 1kHz, matching a sine wave recorded using the exact same setup at SPL of 94 dB. In case of the calibration signal playback level set to match SPL of 94 dB at the eardrum, all the 27 samples should be reproduced at realistic loudness. More details on OCV calibration procedure and other options you can find in Lam et. al. (<a href="https://arxiv.org/abs/2207.12899">2022</a>) and the attached documentation. PLEASE DO NOT EXPOSE YOURSELF NOR THE PARTICIPANTS TO THE CALIBRATION SIGNAL SET AT THE REALISTIC LEVEL AS IT CAN CAUSE HARM.</p> <p><strong>License and reuse</strong>: All SATP recordings are provided under the Creative Commons Attribution 4.0 International (CC BY 4.0) License and are free to use. We encourage other researchers to replicate the SATP protocol and contribute new languages to the dataset. We also encourage the use of these recordings and the perceptual data for further soundscape research purposes. Please provide the proper attribution and get in touch with the authors if you would like to contribute a new translation or for any other collaborations.</p>
Temperature and Climate Attribution estimates supporting "Human Fingerprints on Daily Temperatures in 2022" (2x2 degrees, 2022)
<p>These data support the publication of "Human Fingerprints on Daily Temperatures in 2022" published in the <a href="https://www.ametsoc.org/index.cfm/ams/publications/bulletin-of-the-american-meteorological-society-bams/explaining-extreme-events-from-a-climate-perspective/">BAMS-EEE special issue</a> in 2024 (DOI: <a href="https://doi.org/10.1175/BAMS-D-23-0264.1">10.1175/BAMS-D-23-0264.1</a>). Included are:</p> <ul> <li>Temperatures: <strong>Gilfordetal2024_BAMS-EEE_T2022.nc</strong></li> <li>Attributions estimates (Climate Shift Index and Change in Information due to Perspective): <strong>Gilfordetal2024_BAMS-EEE_ChIP2022.nc</strong></li> </ul> <p>And an accompanying land-sea mask from ERA5 (<strong>Gilfordetal2024_BAMS-EEE_LandSeaMask.nc</strong>). All data values valid for the 2022 calendar year and interpolated to a 2x2 degrees spatial grid to support the study's analysis.</p> <p>For more information on this dataset or to follow up, please contact Daniel Gilford (<a href="mailto:dgilford@climatecentral.org" target="_blank" rel="noopener">dgilford@climatecentral.org</a>).<br><br><em>Funding for this work was provided by the Bezos Earth Fund, The Schmidt Family Foundation, High Meadows Foundation, and the William and Flora Hewlett Foundation.</em></p>
Soil moisture sensor network, design, location attributes and soil properties, Hainich, Germany, project AquaDiva
<p>This dataset contains information of the small scale highly resolved soil moisture measurement network that is part of the of the AquaDiva Critical Zone exploratory, Hainich National Park, Germany. The dataset contains information on soil measurement locations, as well as attributes to the location, the design type (random locations vs transects), as well as locations attributes like distance to the next tree and soil properties. Measurement design was first introduced by Metzger et al., (2017), and used in Fischer et al., 2023. See there for more information.</p> <p><strong>References</strong></p> <p>Fischer-Bedtke, C., Metzger, J. C., Demir, G., Wutzler, T., and Hildebrandt, A.: Throughfall spatial patterns translate into spatial patterns of soil moisture dynamics – empirical evidence, Hydrology and Earth System Sciences, https://doi.org/10.5194/hess-2022-418, 2023.</p> <p>Metzger, J. C., Wutzler, T., Dalla Valle, N., Filipzik, J., Grauer, C., Lehmann, R., Roggenbuck, M., Schelhorn, D., Weckmüller, J., Küsel, K., Totsche, K. U., Trumbore, S., and Hildebrandt, A.: Vegetation impacts soil water content patterns by shaping canopy water fluxes and soil properties, Hydrological Processes, 31, 3783–3795, https://doi.org/10.1002/hyp.11274, 2017.</p>
Water quality and watershed attributes of 41 Pampean streams in Argentina, 12 years later (2003-2015).
This database consists of water chemistry (pH, conductivity, dissolved oxygen, nutrients, and carbonates) and catchment attributes for 41 streams of Buenos Aires province, Argentina. Water quality was measured in 2003/4 and 12 years later (2015/16). Sampling were made in May (autumn), November (spring), and February (summer) at baseflow condition. Some physico-chemical parameters were measured in situ. Parameters determined at laboratory were nutrients and salts. And catchment attributes were determined (physiographic parameters, land use, soil type and geology).
IPBES Data Management Tutorials - Session 5.2: Tools to find and attribute DOIs
<p>The <em>IPBES data management tutorials</em> are short videos to help experts implement the IPBES data management Policy. They cover topics ranging from data management policy, reports, active research data, tools, and examples.</p> <p>The<em> Tools for data management </em>chapter provides IPBES authors with an overview of open source tools used frequently by the scientific community to help it implement data management for the entire data life cycle.</p> <p>The session on <em>tools to find and attribute DOIs </em>covers fundamental background information on digital object identifiers and how to resolve and reserve them.</p>
Parish church (Deerlijk, Parochiekerk Sint-Columba). Altarpiece. Attributed to Jan Demeyere. Ca 1535. 3
<u>File Name</u>: PM_142890_B_Deerlijk <br><u>Sublocation</u>: Parochiekerk Sint-Columba <br><u>Location</u>: Deerlijk <br><u>Province</u>: West-Vlaandren <br><u>Country</u>: Belgium <br><u>Header</u>: Retabel van de Heilige Columba van Sens, toegeschreven aan Jan Demeyere, ca 1535, detail, de doop van de Heilige Columba <br><u>Description</u>: Parish church (Deerlijk, Parochiekerk Sint-Columba). Altarpiece. Attributed to Jan Demeyere. Ca 1535. <br><u>Author</u>: Photo: Paul M.R. Maeyaert <br><u>Author Mail</u>: pmrmeaeyaert@gmail.com <br><u>Copyright</u>: © Paul M.R. Maeyaert,pmrmaeyaert@gmail.com <br><u>Keywords</u>: Europe|Belgium|West-Vlaanderen; Europe|Belgium; Europe|Belgium|West-Vlaanderen|Deerlijk; Cultural heritage|Techniques|Sculpture; Cultural heritage|Techniques; Cultural heritage; Europe|Belgium|Vlaanderen streken|Zuidwest <br><u>Date of Generation</u>: 2022-02-28T14:25:12.044+02:00
Parish church (Deerlijk, Parochiekerk Sint-Columba). Altarpiece. Attributed to Jan Demeyere. Ca 1535. 2
<u>File Name</u>: PM_142892_B_Deerlijk <br><u>Sublocation</u>: Parochiekerk Sint-Columba <br><u>Location</u>: Deerlijk <br><u>Province</u>: West-Vlaandren <br><u>Country</u>: Belgium <br><u>Header</u>: Retabel van de Heilige Columba van Sens, toegeschreven aan Jan Demeyere, ca 1535, detail, tafereel 2: de Heilige Columba wordt voor keizer Aurelianus geleid <br><u>Description</u>: Parish church (Deerlijk, Parochiekerk Sint-Columba). Altarpiece. Attributed to Jan Demeyere. Ca 1535. <br><u>Author</u>: Photo: Paul M.R. Maeyaert <br><u>Author Mail</u>: pmrmeaeyaert@gmail.com <br><u>Copyright</u>: © Paul M.R. Maeyaert,pmrmaeyaert@gmail.com <br><u>Keywords</u>: Europe|Belgium|West-Vlaanderen; Europe|Belgium; Europe|Belgium|West-Vlaanderen|Deerlijk; Cultural heritage|Techniques|Sculpture; Cultural heritage|Techniques; Cultural heritage; Europe|Belgium|Vlaanderen streken|Zuidwest <br><u>Date of Generation</u>: 2022-02-28T14:23:33+02:00
Parish church (Deerlijk, Parochiekerk Sint-Columba). Altarpiece. Attributed to Jan Demeyere. Ca 1535.
<u>File Name</u>: PM_142888_B_Deerlijk <br><u>Sublocation</u>: Parochiekerk Sint-Columba <br><u>Location</u>: Deerlijk <br><u>Province</u>: West-Vlaandren <br><u>Country</u>: Belgium <br><u>Header</u>: Retabel van de Heilige Columba van Sens, toegeschreven aan Jan Demeyere, ca 1535 <br><u>Description</u>: Parish church (Deerlijk, Parochiekerk Sint-Columba). Altarpiece. Attributed to Jan Demeyere. Ca 1535. <br><u>Author</u>: Photo: Paul M.R. Maeyaert <br><u>Author Mail</u>: pmrmeaeyaert@gmail.com <br><u>Copyright</u>: © Paul M.R. Maeyaert,pmrmaeyaert@gmail.com <br><u>Keywords</u>: Europe|Belgium|West-Vlaanderen; Europe|Belgium; Europe|Belgium|West-Vlaanderen|Deerlijk; Cultural heritage|Techniques|Sculpture; Cultural heritage|Techniques; Cultural heritage; Europe|Belgium|Vlaanderen streken|Zuidwest <br><u>Date of Generation</u>: 2022-02-28T13:29:01+02:00
Water levels at tide gauges from: Reconstruction of hourly coastal water levels and counterfactuals without sea level rise for impact attribution
<p>Data to reproduce the analysis of the Hourly Coastal water levels with Counterfactual (HCC) dataset, presented in the publication "<strong>Reconstruction of hourly coastal water levels and counterfactuals without sea level rise for impact attribution</strong>" published in Earth System Science Data (ESSD). </p><p>Note that in this repository, water levels are only provided tide gauge locations which were used for the analysis presented in the paper. The full Hourly Coastal water levels with Counterfactual (HCC) dataset is published in the <a href="https://doi.org/10.48364/ISIMIP.749905">ISIMIP repository</a>.</p><h2>File Descriptions</h2><h4>HCC_analysis_and_plots.ipynb</h4><p>This jupyter-notebook contains all scripts to produce the plots presented in the paper. Make sure that all necessary python packages are installed. The script assumes all netCDF files from this repository to be stored in a sub-directory called "data".</p><h3>hcc_gesla3_99pctl_surge_2011_2015.nc</h3><p>Extreme surge levels from 2011-2015 at 999 GESLA-3 tide gauge stations with at least 90 percent of data in the considered period. As astronomical tides are removed from the modeled and observed water levels to yield the surge component. The file also contains monthly relative water levels and monthly geocentric water levels from 1900-2015 from the HCC dataset.</p><h4>Variables:</h4><ul><li><i>observed_99pctl_surge_level_anomaly</i> -- 99th percentile of daily maximum surge level anomalies from 2011-2015</li><li><i>hcc_99pctl_surge_level_anomaly -- </i>HCC surge level anomalies at the same time steps as <i>observed_99pctl_surge_level_anomaly</i></li><li><i>hcc_counterfactual_99pctl_surge_level_anomaly</i> -- HCC counterfactual surge levels at the same time steps as <i>observed_99pctl_surge_level_anomaly</i></li><li><i>hcc_water_level_monthly</i> – Monthly relative water level from 1900-2015</li><li><i>hcc_geocentric_water_level_monthly</i> – Monthly geocentric water level from 1900-2015</li></ul><h3>hcc_hr_psmsl_water_level_monthly_1900_2015.nc</h3><p>Monthly water levels at 663 PSMSL tide gauge stations of at least 20 year length and with at least 30 percent data coverage in the 1993-2012 period. The file contains data from the HCC, HR and PSMSL datasets. To align PSMSL and HR with HCC, the 1993-2012 average from PSMSL and HR is removed from each of those datasets respectively and the 1993-2012 average of HCC is added. The average is calculated only over all time steps where the associated observational record has valid data.</p><h4>Variables:</h4><ul><li><i>hcc_water_level_monthly</i> – Monthly relative water level from the HCC dataset</li><li><i>hr_aligned_water_level_monthly</i> -- Monthly relative water level from the HR dataset, aligned with <i>hcc_water_level_monthly</i></li><li><i>psmsl_aligned_water_level_monthly</i> -- Monthly relative water level from the PSMSL database, aligned with <i>hcc_water_level_monthly</i></li></ul><h3>hcc_codec_hr_gesla3_water_level_hourly_monthly_1979_2015.nc</h3><p>Hourly water levels at 1040 GESLA-3 tide gauge stations which have at least 30 percent of valid observations between 1979 and 2015. The file contains data from the HCC, CoDEC, HR and GESLA-3 datasets. The different records are not vertically aligned.</p><h4>Variables:</h4><ul><li><i>gesla3_water_level_hourly</i> -- Hourly relative water level from the GESLA3 database</li><li><i>hcc_water_level_hourly</i> -- Hourly relative water level from the HCC dataset</li><li><i>codec_water_level_hourly</i> -- Hourly relative water level from the CoDEC dataset</li><li><i>hr_water_level_monthly</i> -- Monthly relative water level from the HR dataset</li></ul><h3> </h3><h3>hcc_gesla3_water_level_hourly_2011_2015.nc</h3><p>Water levels from the HCC and GESLA-3 datasets, only for tide gauge stations with a complete record in the period 2011-2015 and associated HCC grid points.</p><h4>Variables:</h4><ul><li><i>gesla3_water_level_hourly</i> -- Hourly relative water level from the GESLA3 database</li><li><i>hcc_water_level_hourly</i> -- Hourly relative water level from the HCC dataset</li></ul><h3>slr_ds_psmsl_selected.nc</h3><p>Linear estimates of relative sea level rise from 1900 to 2015. Data is provided at 663 PSMSL tide gauge stations of at least 20 year length and with at least 30 percent data coverage in the 1993-2012 period. Estimates are calculated for the HCC, HR and PSMSL datasets.</p><h4>Variables:</h4><ul><li><i>psmsl_rslr, psmsl_rslr_lower, psmsl_rslr_upper</i> -- Relative sea level rise for PSMSL with lower and upper bounds for a 95 percent confidence interval</li><li><i>hcc_long_rslr, hcc_long_rslr_lower, hcc_long_rslr_upper </i>-- Relative sea level rise for HCC with lower and upper bounds for a 95 percent confidence interval</li><li><i>hr_rslr, hr_rslr_lower, hr_rslr_upper</i> -- Relative sea level rise for HR with lower and upper bounds for a 95 percent confidence interval</li></ul><h3>reg_mask_xr.nc</h3><p>Split of the world into 7 ocean basins: Indian Ocean - South Pacific, Northwest Pacific, East Pacific, South Atlantic, Subtropical North Atlantic, Subpolar North Atlantic West and Subpolar North Atlantic East.</p><h4>Variables:</h4><p><i>reg_mask</i> – Float value, representing the ocean basins</p><p> </p>
Attributing decadal climate variability in coastal sea-level trends
<p>The data produced from analysis to be published in Ocean Science Discussions, paper entitled "Attributing decadal climate variability in coastal sea-level trends". NetCDF contains the following sets of fields:</p> <p>1. Indexing: An <em>index</em> and location (<em>lat, lon</em>) of the coastal grid cells, a locator index attributing each cell to Atlantic, Pacific and Indian Ocean basin, a <em>time</em> (decimal year) index.</p> <p>2. NEMO model trends (<em>nemo_<component>_trend</em>): Rolling decadal trends at each coastal grid cell from the NEMO model run for steric, manometric (dynamic) and GRD. The sum of these components gives the equivalent to absolute sea level trend. </p> <p>3. Climate and oceanographic mode indices: The rolling decadal trends in climate indices and the AMOC index calculated from the AMOC model (<em>ci_trend</em>) and their names (<em>ci_index</em>).</p> <p>4. Empirical Orthogonal Function spatial pattern (<em>eof_<basin>_<component>_D</em>) and Principal Component time series (<em>eof_<basin>_<component>_PC</em>)<em> </em>of the NEMO model trends.</p> <p>5. Coefficient of linear regression between PC and climate indices (<em>recon_<basin>_<component>_beta</em>) and the rolling trend time series at each grid cell from the reconstruction, sum{ci_trend*beta} (<em>recon_<basin>_<component>_trend</em>).</p> <p>In 4 and 5, the indices are given by basin. The total coastline is a concatenation of the Atlantic, Pacific and Indian basin data in that order. The absolute SSH is given by the sum of components. i.e. the SSH for all coastal cells in order <em>index</em>:</p> <p>recon_sum_trend([index(Atlantic_index); index(Pacific_index); index(Indian_index)] = ...</p> <p> [recon_Atlantic_manometric_trend+recon_Atlantic_steric_trend+recon_Atlantic_grd_trend; ...</p> <p> recon_Pacific_manometric_trend+recon_Pacific_steric_trend+recon_Pacific_grd_trend; ...</p> <p> recon_Indian_manometric_trend+recon_Indian_steric_trend+recon_Indian_grd_trend]</p>
Self-Attribution of Distorted Reaching Movements in Immersive Virtual Reality Dataset
<p>This dataset accompanies the paper “Self-Attribution of Distorted Reaching Movements in Immersive Virtual Reality” published in the Computer and Graphics journal from Elsevier. It contains 3 datasets related to the experiments described in the paper. All datasets are in “.csv” format and can be easily loaded by statistical analysis tools (e.g. a dataset can be loaded in r using the command read.csv(“filename.csv”)). It also contains the C# Unity implementation of the distortion function presented in the paper.</p> <p>Paper reference:</p> <p>Galvan Debarba H, Boulic R, Salomon R, Blanke O, Herbelin B. Self-Attribution of Distorted Reaching Movements in Immersive Virtual Reality. Computers & Graphics. 2018; ISSN 0097-8493. Elsevier.</p> <p>DOI: doi.org/10.1016/j.cag.2018.09.001</p>
Attributes: A Curriculum Analytics System for measuring learning outcomes - Overview
<p><span><strong>Link to video </strong><a href="https://vimeo.com/1015456231?share=copy#t=0"><strong>https://vimeo.com/1015456231?share=copy - t=0</strong></a><br><br>The Curriculum Analytics System at the Instituto Tecnológico de Costa Rica, integrated into TEC Digital, supports faculty, coordinators, and students in assessing engineering learning outcomes during accreditation processes. The system offers two key user modules: one for coordinators to map and manage learning outcomes, and another for instructors to conduct assessments through the course portal. Coordinators oversee course and attribute mapping using visual representations of study plans, control points, and outcome visualizations. Instructors configure assignments and evaluate student submissions with standardized rating scales. The system tracks progress in real-time and generates</span> <span>comprehensive reports with performance metrics, facilitating continuous improvement in academic programs.</span></p> <p><strong><span>Key words: </span></strong><span>attributes, learning outcomes, TEC Digital, curriculum analytics, continuous improvement. </span></p>
Public attributions made on Bionomia, December 2022
<p>Public attributions of natural history specimens made for collectors and determiners on Bionomia (<a href="https://bionomia.net/">https://bionomia.net</a>) using data from the Global Biodiversity Information Facility, <a href="https://gbif.org/">https://gbif.org</a>. This file was downloaded from the website on the 28 December 2022. More recent versions of the resource can be found on the official repository on Zenodo for Bionomia at <a href="https://doi.org/10.5281/zenodo.13937806" target="_blank" rel="noopener">https://doi.org/10.5281/zenodo.13937806</a> .</p> <p>Each row of the file has a subject, predicate and object...</p> <ul> <li>Subject is the specimen identifer from the Global Biodiversity Information Facility e.g. https://gbif.org/occurrence/1801358422</li> <li>Predicate is the indicator as to whether the action of the person is as an identifier or collector. e.g. http://rs.tdwg.org/dwc/iri/identifiedBy</li> <li>Object Is the attributed preson identifed by their ORCID identifier, or Wikidata identifier. e.g. https://orcid.org/0000-0001-9008-0611</li> </ul>
Fire Self-Limitation (FiSL) Experiment: Quantifying Wildfire Carbon Combustion Losses in boreal Deciduous and Mixed Forests in Interior Alaska and the Boreal Cordillera I: Site Attribute Data 2022
This dataset contains site characteristics collected in the field for plots in 8 fire scars in Interior Alaska and the Yukon. Data was collected in the summer of 2022. Fire scars sampled included Shovel Creek (2019), Aggie Creek (2015), Hess Creek (2019), Baker (2015), Munson Creek (2021), Isom Creek (2020), 2019MA014 (2019), and 2019BC005 (2019). Data includes detailed site characteristics collected at the site level. Each site included three 10 m * 2 m plots (A, B, and C) laid in a single 30 m transect (or, where constrained, in parallel).
Undirected Node Attributed Social Network Graph of Twitter Users interested in plastic pollution - created in the framework of the PlasticTwist project
<p>This dataset has been created in the framework of the Plastic Twist project (<a href="https://ptwist.eu/">Ptwist</a>) and more specifically using the Ptwist crowdsourcing application (<a href="https://crowdsourcing.plastictwist.com/">crowdsourcing.plastictwist.com/</a>). We are sharing the edge list and specific node attributes (hashtags) of Twitter users posting about plastic pollution. The dataset can be used for community detection,clustering, node importance, influence maximization tasks, etc. Each user is represented by a unique integer which has nothing to do with the official Twitter user ID. The dataset contains three (3) files: </p> <ul> <li>ptwist.edgelist: A list containing all the 1,362,863 edges between the users. When loaded they create an undirected graph of 800K+ users.</li> <li>node_attributes.txt: This file contains information about the hashtags used by each user. (e.g. "652003": ["SingleUsePlastic"] -> user 6529003 has used the hashtag SingleUsePlastic) </li> <li>annotated_graph: A pickle file which, when loaded, returns a <a href="https://networkx.github.io/">NetworkX</a> node attributed undirected graph.</li> </ul> <p> </p> <p> </p>
Graphs and Attributes used for the attribute-structure correlation pattern mining
<p>## SCPM: An implementation of an algorithm for structural correlation pattern mining.</p> <p>The structural correlation measures how a set of attributes induces dense subgraphs in an attributed graph. A structural correlation pattern is a dense subgraph induced by a particular attribute set. Structural correlation pattern mining is useful to analyze how different attribute sets are correlated to dense subgraphs in several real-life attributed graphs.</p> <p>**Relevant Publications**</p> <p>* Arlei Silva, Wagner Meira, Jr., and Mohammed J. Zaki. Structural correlation pattern mining for large graphs. In Proceedings of the Eighth Workshop on Mining and Learning with Graphs (MLG '10).</p> <p>* Arlei Silva, Wagner Meira, Jr., and Mohammed J. Zaki. Mining Attribute-structure Correlated Patterns in Large Attributed Graphs. In Proceedings of the VLDB Endowment (PVLDB '12).</p> <p>* Arlei Silva. Structural correlation pattern mining for large graphs. M.Sc Thesis, Computer Science Department, Universidade Federal de Minas Gerais, 2011.</p> <p>* Arlei Silva, Wagner Meira Jr. Structural correlation pattern mining for large graphs. Thesis and Dissertation Contest of the Brazilian Computer Society (CTD'12).</p> <p><br> ## HOW TO</p> <p>cd to trunk and run make<br> see README in trunk</p> <p><br> ## Datasets:</p> <p>### Description:</p> <p>#### ATTRIBUTE FILE:</p> <p>Format: Lists the attributes of each vertex from the graph.</p> <p> <VERTEX_ID>,<ATTRIBUTE_ID>,<ATTRIBUTE_ID>...,<ATTRIBUTE_ID></p> <p> Example: <br> 1,A,C <br> 2,A <br> 3,A,C,D <br> 4,A,D <br> 5,A,E <br> 6,A,B,C <br> 7,A,B,E <br> 8,A,B <br> 9,A,B <br> 10,A,B,D <br> 11,A,B</p> <p>#### GRAPH FILE:</p> <p>Format: Lists the neighbors of each vertex from the graph (adjacency list). Although the graph is undirected, each edge must be included in both directions.</p> <p> <VERTEX_ID>,<NEIGHBOR_ID>,<NEIGHBOR_ID>...,<NEIGHBOR_ID> <br> <br> Example: <br> 1,4 <br> 2,3 <br> 3,2,4,5,6,7 <br> 4,1,3,5,6 <br> 5,3,4,6 <br> 6,3,4,5,7,8,9,10 <br> 7,3,6,8,11 <br> 8,6,7,9,10,11 <br> 9,6,8,10,11 <br> 10,6,8,9,11 <br> 11,7,8,9,10</p> <p>### REAL DATASETS</p> <p>Lastfm:</p> <p>attributes: attrLastFm.csv.tar.bz2</p> <p>network: graphLastFm.csv.tar.gz</p> <p>DBLP:</p> <p>attributes: newAttrDBLP.csv.tar.bz2</p> <p>network: newGraphDBLP.csv.tar.bz2</p> <p>CITESEER:</p> <p>attributes: attrCiteseer.csv.tar.bz2</p> <p>network: graphCiteseer.csv.tar.bz2</p>
Relations in the Biographical Dictionary of Republican China - Node & Edge lists and Attribute File
<p>This dataset contains the node and edge lists of relations in the BDRC. They are based on the "Relations in the Biographical Dictionary of Republican China - Standardized output" file in this repository. The attributes of the nodes refer only to the main 589 figures in the BDRC.</p>
Daily Temperature Attribution
<p>This repository contains the data for Pershing et al. "High-resolution attribution of the daily exposure of people and ecosystems to climate-driven heat" submitted to PNAS. </p><p>The files Temperature_ChIP_<i>{tvar}_2023</i>.nc (tvar = 'tmax','tmin' or 'tavg') contain the daily temperatures, temperature anomalies (relative to 1991-2023) from ERA5. The variable "change_in_information" (ChIP) indicates the influence of climate change on each temperature. ChIP = log2(likelihood of T in the current climate/likelihood on climate without global warming). </p><p>The three .csv files contain daily ChIP data from 1970-2023 aggregated over the land surface or across land biomes using area-weighted averaging and over major countries using population weighting.</p><p>The two files labeled "minimum mortality temperature" contain the annual counts of days above the 84th percentile minimum mortality temperature (or below the 16th percentile temperature). Days that meet the threshold and have ChIP>=1 are also indicated.</p>
Large-scale attributed graph & hypergraph datasets: TWeibo, Amazon2M, Amazon, MAG-PM
<p>Here we provide additional large-scale datasets used in our work "A Versatile Framework for Attributed Network Clustering via K-Nearest Neighbor Augmentation", along with the index files for constructing KNN graphs using ScaNN and Faiss.</p> <p>Usage:</p> <p>cd ANCKA/</p> <p>unzip ~/Download_path/ANCKA_data.zip -d data/</p>
Maps of the detailed spatially and temporally attributed emission for area of Legerova and Sokolska (TURBAN-D18)
<h3>Basic information</h3> <p>This dataset contains six folders with maps of input data for simulations published in project TURBAN as result D17 (see <a href="../records/10982836">https://zenodo.org/records/10982836</a>). Each folder contains air quality inputs for the so-called Legerova domain, an area in the city of Prague, Czech Republic, centred around the traffic-heavy streets Legerova and Sokolská. All times are in UTC (local time in winter, CET, is UTC +01:00, summer time, CEST, is UTC +02:00). In total 6 episodes in 2022 and 2023 were selected:</p> <ol> <li>s1 2022-07-17 00:00:00 - 2022-07-20 00:00:00</li> <li>s2: 2022-08-02 00:00:00 - 2022-08-05 00:00:00</li> <li>s3: 2022-09-22 00:00:00 - 2022-09-25 00:00:00</li> <li>s4: 2022-12-08 00:00:00 - 2022-12-11 00:00:00</li> <li>s5: 2023-01-27 00:00:00 - 2023-01-30 00:00:00</li> <li>s6: 2023-02-13 00:00:00 - 2023-02-16 00:00:00</li> </ol> <p>For more detailed description of the experiments see the <strong>TURBAN</strong> project website at <a href="https://www.project-turban.eu/">https://www.project-turban.eu/</a>.</p> <h3>General organisation, variables and file nomenclature</h3> <p>Each selected epizode (s1-s6) has three subfolders; input files in ASCII (<em>output-ascii</em>) or GeoTiff (<em>output-gis</em>) formats that can be viewed in many GIS applications. In the third subfolder are maps in the PNG format (<em>output-png</em>).</p> <p>Each subfolder includes 4 subfolders with emissions summarized in all layers above ground. Variable <em>vsrc_PM10</em> is the concentration of volume source emissions (VSRC) of the PM10, <em>vsrc_PM25</em> is the concentration of PM2.5, <em>vsrc_NO</em> is the concentration of NO and <em>vsrc_NO2</em> is the concentration of NO2.</p> <p>Each file (PRJ, TIF, ASC or PNG) has the same nomenclature. An example (vsrc_NO_abs-01h_20220717_1200-1300.png) could be parsed as: variable name (vsrc_NO), processed input (abs-01h), date (20220717) and period (1200-1300). So, the result is a map with emission fluxes of NO between 12:00 and 13:00 UTC 24 Jul 2019.</p> <h3>Emissions (see section 2.4.3 in Resler et al., 2024)</h3> <p>The data were processed from datasets published by CHMI, data collected by the Municipality of Prague and its organizations, data obtained by the researcher (ATEM) while providing expert studies in the past, and results of previous research projects. The input data of the used emission sources can be divided into two basic groups: emission from local heating and transport sources.</p> <p>Emissions for local heating were determined by calculations based on data from CHMI and the Czech Statistical Office (CZSO). Emissions from the transport sources were modeled using the MEFA transportation emission model which is recommended for the use in the Czech Republic by the Ministry of Environment of the Czech Republic. The model takes into account factors such as road gradient, the number of vehicles on the road, the flow of traffic, the composition of car types, and the emission characteristics of the individual car types. The emission calculation is based on data from the traffic census provided by the Prague Technical Administration of Roads (TSK Praha) and on data from the census of the composition of the transportation fleet in Prague built in the MEFA emission model. The data are based on regular surveys of the fleet composition carried out in Prague (Karel et al., 2021). The dust resuspension was computed according to the methodology published by the Ministry of Environment (Karel et. al., 2015). This methodology is based on US EPA methodology AP-42 (EPA, 2011) and was adjusted for the conditions of the Czech Republic. For the garages and parking lots, the results of the project TH03030496 (Karel et al., 2020) were used and for the bus stations, publicly available data about transportation were gathered from the Prague Public Transit Company (DPP).</p> <p>The disaggregation of the annual emissions into hourly intervals was then performed according to the type of source. For combustion sources distribution of emissions to days was done according to natural gas supply profiles for category DOM4 were used (OTE, 2024) and complemented by daily profiles for SNAP 2 (van der Gon, 2011). For transport sources, the census data from TSK Praha was utilized for all streets where it was available. For Legerova and Sokolská streets, hourly traffic intensity data were obtained and used directly for the selected episodes. For streets that were not covered by regular traffic surveys, the spatial and temporal distribution of the traffic intensities were based on analysis and evaluation of the relevant studies for the particular area (e.g. urban planning studies, Environmental Impact Assessment (EIA), etc.) and combined with information like street type, location, traffic regime, and pavement type. This approach allowed us to specify the distribution of the transportation intensities on smaller streets. For the detailed modeling of emissions from rail transport (diesel locomotives), the data of train rides were obtained from the Railway Administration (SŽ) and emission factors from the EMEP/EEA Air Pollutant Emission Inventory Guidebook 2019 (EEA, 2019) were used. Emissions from river ships were obtained from the CHMI national database and spatially distributed to the area of the river.</p> <p>Spatial transformation of the line and point emission into the corresponding areas was done with the utilization of the surrogates representing corresponding areas (e.g. areas of the street traffic lines and parking places for traffic emission and areas of the building roofs for local heating sources). This not only ensured the reasonable spatial distribution of the emission in the street canyon but also decreased the gradients of the emission field and with this proneness of the model to numerical inaccuracy of the micro-scale model. The processing of the emission sources into hourly emission flows was done in the emission model FUME recently extended for processing of the PALM emission (Belda et al., 2024).</p> <h3>Acknowledgements</h3> <p>The PALM simulations, and pre- and postprocessing were performed partially on the HPC infrastructure of the Institute of Computer Science of the Czech Academy of Sciences (ICS), supported by the long-term strategic development financing of the ICS (RVO:67985807) and partially on the IT4I HPC infrastructure supported by the Ministry of Education, Youth and Sports of the Czech Republic through the e-INFRA CZ (ID:90254). The work was performed within the project TURBAN (TO01000219; TURBAN – Turbulent-resolving urban modelling of air quality and thermal comfort) supported by Norway Grants and Technology Agency of the Czech Republic.</p> <h3>Literature</h3> <p>Note that some sources are available only in Czech language.</p> <p>Belda, M., et al. (2024) FUME 2.0 – Flexible Universal processor for Modeling Emissions, EGUsphere [preprint]. <a href="https://doi.org/10.5194/egusphere-2023-2740">https://doi.org/10.5194/egusphere-2023-2740</a></p> <p>Karel, J., et al. (2020) Projekt TH03030496 - Zmapování a emisní bilance neevidovaných zdrojů emisí znečišťujících látek na území městských aglomerací. Mapa neevidovaných zdrojů emisí znečišťujících látek na území aglomerace CZ01 Praha. Partially available at: <a href="https://www.atem.cz/neevidovane_zdroje.php">https://www.atem.cz/neevidovane_zdroje.php</a></p> <p>Karel, J., et al. (2015) Metodika pro výpočet emisí částic pocházejících z resuspenze ze silniční dopravy, CENEST, s. r. o., Prague. Available at: <a href="https://www.mzp.cz/C1257458002F0DC7/cz/doprava/$FILE/OOO-resuspenze_metodika-20190708.pdf">https://www.mzp.cz/C1257458002F0DC7/cz/doprava/$FILE/OOO-resuspenze_metodika-20190708.pdf</a></p> <p>Karel J., et. al. (2021) Zpráva o dynamické skladbě vozového parku na území hlavního města Prahy v roce 2020, Prague 2021. Available upon request from the Environmental Protection Division of the Prague Municipality.</p> <p>EPA (2011) Compilation of Air Pollutant Emission Factors, Volume I, AP-42. Section 13.2.1. Paved roads. EPA Research Triangle Park, US, 2003, updated 2011. Available at: <a href="https://www.epa.gov/air-emissions-factors-and-quantification/ap-42-compilation-air-emissions-factors-stationary-sources">https://www.epa.gov/air-emissions-factors-and-quantification/ap-42-compilation-air-emissions-factors-stationary-sources</a></p> <p>van der Gon, H.D., et al. (2011) Description of Current Temporal Emission Patterns and Sensitivity of Predicted AQ for Temporal Emission Patterns. EU FP7 MACC Deliverable Report D_D-EMIS_1.3. Available at: <a href="https://atmosphere.copernicus.eu/sites/default/files/2019-07/MACC_TNO_del_1_3_v2.pdf">https://atmosphere.copernicus.eu/sites/default/files/2019-07/MACC_TNO_del_1_3_v2.pdf</a></p> <p>EEA (2019) European Environment Agency, EMEP/EEA air pollutant emission inventory guidebook 2019 – Technical guidance to prepare national emission inventories, Publications Office. Available at: <a href="https://data.europa.eu/doi/10.2800/293657">https://data.europa.eu/doi/10.2800/293657</a></p> <p>OTE (2024) Gas Load Profiles - temperature and recalculated TDD. Available at: <a href="https://www.ote-cr.cz/en/statistics/gas-load-profiles/normalized-lp?set_language=en">https://www.ote-cr.cz/en/statistics/gas-load-profiles/normalized-lp?set_language=en</a></p> <p> </p> <p> </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.