Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

6,250

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

6,250 results for “classification”

Learn how ShareScore rates datasets ↗
edi60/100

LAGOS-US HUMAN v2: Data module of human population(1990-2020), urbanization classification, and lake access in the conterminous U.S.

The LAGOS-US HUMAN v1 data package is an extension module of the LAGOS-US research platform that includes data characterizing human population (population count, race, ethnicity, socioeconomic information), urbanization, and lake access of 479,950 lakes larger than or equal to 1 ha in the conterminous U.S. (48 states plus the District of Columbia). This data module contains four data tables linked through the unique lake identifier for the LAGOS-US research platform, lagoslakeid. Human population characteristics (race, ethnicity, and socioeconomic factors) were derived from U.S. census data for 1990, 2000, 2010, and 2020. Lakes were classified as urban or not using two different classifications: one based on the ‘Developed’ land category in the National Land Cover Dataset; and another based on the 2020 Census Urban Areas category. Metrics for lake access were developed from national datasets on public boat launches, transportation, and public lands. LAGOS-US HUMAN v1 provides a link between lake data and human contexts, facilitating interdisciplinary research in limnology, urban ecology, environmental justice, and conservation. To facilitate such studies, users are encouraged to use the other three core data modules of the LAGOS-US platform: LOCUS (location, identifiers, and physical characteristics of lakes and their watersheds); GEO (geospatial ecological context at multiple spatial and temporal scales); and LIMNO (in situ lake physical, chemical, and biological measurements through time) that are each found in their own data packages.

openCC (other)Oct 2025View details →
edi60/100

Wrack classification data based on UAV imagery from Dean Creek on Sapelo Island, GA

We used a DJI Matrice 210 UAV with a MicaSense Altum to collect a total of 20 images from January 2020 - December 2021 in a the Dean Creek marsh on Sapelo Island, GA. Wrack was classified using a principal component analysis. Wrack patches under 1 m2 were excluded from analyses. Wrack classifications were converted to polygon and point data where each point represents a 5 cm x 5 cm pixel. Those files were then used to analyze wrack characteristics, their relation to environmental drivers, and landscape based patterns. For both polygon and point data, we used the National Elevation Dataset (https://gdg.sc.egov.usda.gov/Catalog/ProductDescription/NED.html) to determine the elevation of each wrack patch. Creeks and shorelines were digitized and used to determine each wrack patches' distance to water. We calculated the frequency of wrack deposition at each point by adding together the number of images where that pixel was classified as wrack over the course of the study. Polygon data were related to tide height from a NOAA tidal station data product (Ft. Pulaski, Station 8670870; https://tidesandcurrents.noaa.gov) and wind speed and wind direction from the Marsh Landing weather station (downloaded data for the SAPMLMET met station from: https://cdmo.baruch.sc.edu/) to evaluate the relationship of wrack to environmental drivers.

openCC (other)Aug 2023View details →
zenodo52/100

Global Extra-tropical Circulation Database based on the Jenkinson-Collison Classification calculated with 6-hourly mean sea-level pressure fields from various reanalysis datasets

<h1>Dataset Description</h1> <p>Global Extra-tropical Circulation Database based on the Jenkinson-Collison Classification calculated with 6-hourly mean sea-level pressure fields from several reanalysis datasets. This dataset is the result of an extension of the Jenkinson-Collison circulation type classification to the entire globe, including a modification of its original formulation for the southern hemisphere.</p> <p>A modified version of the IPCC-AR6 Reference Regions that excludes the intertropical range where the method is not applicable is also included, as used in the reference paper for global assessment.</p> <p>Further details in <a href="https://doi.org/10.1007/s00382-022-06658-7" target="_blank" rel="noopener">https://doi.org/10.1007/s00382-022-06658-7&nbsp;</a></p> <h2>Note for version 1.1.0</h2> <p>This version corrects an issue in the previous release, which was incorrectly labeled as <em>version 0.1</em>. That version was incomplete due to the omission of previously existing files, and should be considered <strong>incomplete</strong>. Version 1.1.0 restores all original files alongside the newly added one, ensuring the dataset is now complete and consistent. We apologize for any inconvenience this may have caused and appreciate your understanding.</p>

opencc-by-4.0Dec 2021View details →
zenodo52/100

Dataset for: A continuous classification of the 480,000 lakes of the conterminous US based on geographic archetypes

<p>These datasets were used in a journal article with the goal of developing a new geographic classification approach for ~480,000 lakes ≥ 1 ha in the conterminous U.S. based on archetypes defined as endmembers&nbsp;with distinct combinations of climate, hydrologic, geologic, topographic, and morphometric properties. We identified seven lake archetypes; each study lake was then assigned weights for each of the archetypes. The data used to develop the archetypes, archetype weights, and variables used in associated analyses is provided in three data tables. The first includes the lake-specific transformed predictors used to generate the seven archetypes, the weights corresponding to each archetype, the archetype with the maximum weight and the weight of that maximum archetype. The second provides lake-specific raw values for each predictor and for the 19 response variables used to explore aspects of the archetype classification. The final metadata table provides a data dictionary for all columns in the previously mentioned data tables.</p>

opencc-by-4.0Oct 2023View details →
zenodo52/100

Dataset for training the Surrogate Model of microlaser neurons on the reduced MNIST classification task

<p>This dataset was used to train a surrogate multilayer perceptron surrogate model of microlaser neurons.</p> <p>It is in csv format. It was generated using the Yamada Model as found in&nbsp;</p> <p><span>Selmi F, Braive R, Beaudoin G, Sagnes I, Kuszelewicz R and Barbay S 2014 Relative Refractory Period in an Excitable Semiconductor Laser <em>Phys. Rev. Lett.</em> <strong>112</strong> 183902</span>.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Classification of web-based Digital Humanities projects leveraging information visualisation techniques

<h2>Description</h2> <p>This dataset contains a list of 186 Digital Humanities projects leveraging information visualisation methods. Each project has been classified according to visualisation and interaction techniques, narrativity and narrative solutions, domain, methods for the representation of uncertainty and interpretation, and the employment of critical and custom approaches to visually represent humanities data.</p> <p>&nbsp;</p> <h2>Classification schema: categories and columns</h2> <p>The <code>project_id</code> column contains unique internal identifiers assigned to each project. Meanwhile, the&nbsp;<code>last_access</code> column records the most recent date (in DD/MM/YYYY format) on which each project was reviewed based on the web address specified in the <code>url</code> column.<br>The remaining columns can be grouped into descriptive categories aimed at characterising projects according to different aspects:</p> <p>&nbsp;</p> <p><strong>Narrativity.</strong> It reports the presence of information visualisation techniques employed within narrative structures. Here, the term narrative encompasses both author-driven linear data stories and more user-directed experiences where the narrative sequence is determined by user exploration [1]. We define 2 columns to identify projects using visualisation techniques in narrative, or non-narrative sections. Both conditions can be true for projects employing visualisations in both contexts. Columns:</p> <ul> <li> <p><code>non_narrative</code> (boolean)</p> </li> <li> <p><code>narrative</code> (boolean)</p> </li> </ul> <p>&nbsp;</p> <p><strong>Domain.</strong> The humanities domain to which the project is related. We rely on [2] and the chapters of the first part of [3] to abstract a set of general domains. Column:</p> <ul> <li> <p><code>domain</code> (categorical):</p> </li> <ul> <li> <p>History and archaeology</p> </li> <li> <p>Art and art history</p> </li> <li> <p>Language and literature</p> </li> <li> <p>Music and musicology</p> </li> <li> <p>Multimedia and performing arts</p> </li> <li> <p>Philosophy and religion</p> </li> <li> <p>Other: both extra-list domains and cases of collections without a unique or specific thematic focus.</p> </li> </ul> </ul> <p>&nbsp;</p> <p><strong>Visualisation of uncertainty and interpretation.</strong> Buiding upon the frameworks proposed by [4] and [5], a set of categories was identified, highlighting a distinction between precise and impressional communication of uncertainty. Precise methods explicitly represent quantifiable uncertainty such as missing, unknown, or uncertain data, precisely locating and categorising it using visual variables and positioning. Two sub-categories are interactive distinction, when uncertain data is not visually distinguishable from the rest of the data but can be dynamically isolated or included/excluded categorically through interaction techniques (usually filters); and visual distinction, when uncertainty visually &ldquo;emerges&rdquo; from the representation by means of dedicated glyphs and spatial or visual cues and variables. On the other hand, impressional methods communicate the constructed and situated nature of data [6], exposing the interpretative layer of the visualisation and indicating more abstract and unquantifiable uncertainty using graphical aids or interpretative metrics. Two sub-categories are: ambiguation, when the use of graphical expedients&mdash;like permeable glyph boundaries or broken lines&mdash;visually convey the ambiguity of a phenomenon; and interpretative metrics, when expressive, non-scientific, or non-punctual metrics are used to build a visualisation. Column:</p> <ul> <li> <p><code>uncertainty_interpretation</code> (categorical):</p> </li> <ul> <li> <p>Interactive distinction</p> </li> <li> <p>Visual distinction</p> </li> <li> <p>Ambiguation</p> </li> <li> <p>Interpretative metrics</p> </li> </ul> </ul> <p>&nbsp;</p> <p><strong>Critical adaptation.</strong> We identify projects in which, with regards to at least a visualisation, the following criteria are fulfilled: 1) avoid repurposing of prepackaged, generic-use, or ready-made solutions; 2) being tailored and unique to reflect the peculiarities of the phenomena at hand; 3) avoid simplifications to embrace and depict complexity, promoting time-consuming visualisation-based inquiry. Column:</p> <ul> <li> <p><code>critical_adaptation</code> (boolean)</p> </li> </ul> <p>&nbsp;</p> <p><strong>Non-temporal visualisation techniques.</strong> We adopt and partially adapt the terminology and definitions from [7]. A column is defined for each type of visualisation and accounts for its presence within a project, also including stacked layouts and more complex variations. Columns and inclusion criteria:</p> <ul> <li> <p><code>plot</code> (boolean): visual representations that map data points onto a two-dimensional coordinate system.</p> </li> <li> <p><code>cluster_or_set</code> (boolean): sets or cluster-based visualisations used to unveil possible inter-object similarities.</p> </li> <li> <p><code>map</code> (boolean): geographical maps used to show spatial insights. While we do not specify the variants of maps (e.g., pin maps, dot density maps, flow maps, etc.), we make an exception for maps where each data point is represented by another visualisation (e.g., a map where each data point is a pie chart) by accounting for the presence of both in their respective columns.</p> </li> <li> <p><code>network</code> (boolean): visual representations highlighting relational aspects through nodes connected by links or edges.</p> </li> <li> <p><code>hierarchical_diagram</code> (boolean): tree-like structures such as tree diagrams, radial trees, but also dendrograms. They differ from networks for their strictly hierarchical structure and absence of closed connection loops.</p> </li> <li> <p><code>treemap</code> (boolean): still hierarchical, but highlighting quantities expressed by means of area size. It also includes circle packing variants.</p> </li> <li> <p><code>word_cloud</code> (boolean): clouds of words, where each instance&rsquo;s size is proportional to its frequency in a related context</p> </li> <li> <p><code>bars</code> (boolean): includes bar charts, histograms, and variants. It coincides with &ldquo;bar charts&rdquo; in [7] but with a more generic term to refer to all bar-based visualisations.</p> </li> <li> <p><code>line_chart</code> (boolean): the display of information as sequential data points connected by straight-line segments.</p> </li> <li> <p><code>area_chart</code> (boolean): similar to a line chart but with a filled area below the segments. It also includes density plots.</p> </li> <li> <p><code>pie_chart</code> (boolean): circular graphs divided into slices which can also use multi-level solutions.</p> </li> <li> <p><code>plot_3d</code> (boolean): plots that use a third dimension to encode an additional variable.</p> </li> <li> <p><code>proportional_area</code> (boolean): representations used to compare values through area size. Typically, using circle- or square-like shapes.</p> </li> <li> <p><code>other</code> (boolean): it includes all other types of non-temporal visualisations that do not fall into the aforementioned categories.</p> </li> </ul> <p>&nbsp;</p> <p><strong>Temporal visualisations and encodings.</strong> In addition to non-temporal visualisations, a group of techniques to encode temporality is considered in order to enable comparisons with [7]. Columns:</p> <ul> <li> <p><code>timeline</code> (boolean): the display of a list of data points or spans in chronological order. They include timelines working either with a scale or simply displaying events in sequence. As in [7], we also include structured solutions resembling Gantt chart layouts.</p> </li> </ul> <ul> <li> <p><code>temporal_dimension</code> (boolean): to report when time is mapped to any dimension of a visualisation, with the exclusion of timelines. We use the term &ldquo;dimension&rdquo; and not &ldquo;axis&rdquo; as in [7] as more appropriate for radial layouts or more complex representational choices.</p> </li> <li> <p><code>animation</code> (boolean): temporality is perceived through an animation changing the visualisation according to time flow.</p> </li> <li> <p><code>visual_variable</code> (boolean): another visual encoding strategy is used to represent any temporality-related variable (e.g., colour).</p> </li> </ul> <p>&nbsp;</p> <p><strong>Interaction techniques.</strong> A set of categories to assess affordable interaction techniques based on the concept of user intent [8] and user-allowed data actions [9]. The following categories roughly match the &ldquo;processing&rdquo;, &ldquo;mapping&rdquo;, and &ldquo;presentation&rdquo; actions from [9] and the manipulative subset of methods of the &ldquo;how&rdquo; an interaction is performed in the conception of [10]. Only interactions that affect the visual representation or the aspect of data points, symbols, and glyphs are taken into consideration. Columns:</p> <ul> <li> <p><code>basic_selection</code> (boolean): the demarcation of an element either for the duration of the interaction or more permanently until the occurrence of another selection.</p> </li> <li> <p><code>advanced_selection</code> (boolean): the demarcation involves both the selected element and connected elements within the visualisation or leads to brush and link effects across views. Basic selection is tacitly implied.</p> </li> <li> <p><code>navigation</code> (boolean): interactions that allow moving, zooming, panning, rotating, and scrolling the view but only when applied to the visualisation and not to the web page. It also includes &ldquo;drill&rdquo; interactions (to navigate through different levels or portions of data detail, often generating a new view that replaces or accompanies the original) and &ldquo;expand&rdquo; interactions generating new perspectives on data by expanding and collapsing nodes.</p> </li> <li> <p><code>arrangement</code> (boolean): methods to organise visualisation elements (symbols, glyphs, etc.) or multi-visualisation layouts spatially through drag and drop or according to a criterion via more automatic triggers.</p> </li> <li> <p><code>change</code> (boolean): visual encoding alterations involving different aspects of visualisation as a whole: the same content is presented with another visualisation technique; the change involves symbols or glyphs aspect (colour, size, shape, etc.); the visualisation type is unaltered, but the layout variant changes (e.g., to stacked layouts); or other changes like axes inversion and scale modifications. The presence of all the visualisation techniques involved in a change is reported.</p> </li> <li> <p><code>visualisation_filter</code> (boolean): filters to exclude or include visualisation elements with respect to defined criteria, without reloading or generating a new visualisation. Unlike options triggering the fetch of new data to alter the visualisation content, filters seamlessly operate on existing visual elements.</p> </li> <li> <p><code>collection_filter</code> (boolean): the interaction with visualised elements acts as a filter for a related collection or list of items (e.g., clicking a region on a map filters a list of items according to spatial metadata).</p> </li> <li> <p><code>aggregation</code> (boolean): changes to the granularity of visual elements according to a variable. It produces either visual data summarisations or segregations.</p> </li> <li> <p><code>btfw_interaction</code> (boolean): to identify the use of &ldquo;breaking the fourth wall interactions&rdquo; as defined [11]. It applies only to narratives.</p> </li> </ul> <p>&nbsp;</p> <p><strong>Narrative flow factors.</strong> Other categories aim to identify patterns in the design of narrative solutions. It is worth noticing that a project with multiple and diverse narratives can potentially report multiple design choices for the same column. Part of the factors and definitions from [12] are here re-used and adapted.</p> <p><em>Story layout </em>columns define the layout, or genre, of the narrative format:</p> <ul> <li> <p><code>document_layout</code> (boolean)</p> </li> <li> <p><code>slideshow_layout</code> (boolean)</p> </li> <li> <p><code>hybrid_layout</code> (boolean): mixing document and slideshow layouts.</p> </li> <li> <p><code>other_layout</code> (boolean): more complex solutions.</p> </li> </ul> <p><em>Role of visualisation</em> columns describe the role visualisations detain with respect to the entire story, in particular, with reference to the textual part of the narratives:</p> <ul> <li><code>equal_role</code> (boolean): visualisations and text play an equal role in the narrative.</li> <li><code>figure_role</code> (boolean): visualisations are supporting elements compared to the role of text.</li> <li><code>annotated_role</code> (boolean): visualisations are the drivers of the narrative.</li> </ul> <p><em>Story progression</em> columns categorise the shape of possible story paths:</p> <ul> <li> <p><code>linear_progression</code> (categorical): strongly author-driven or user-directed narrative. Possible values specify the potential to skip certain parts while not having a fully explorative experience:</p> </li> <ul> <li> <p>Skip</p> </li> <li> <p>No-skip</p> </li> </ul> <li> <p><code>user_directed</code> (bool): users can select a path among multiple alternatives and compose narrative pieces, providing a broder degree of interaction and exploration possibilities [1]. If a linear path can be suggested, here it remains merely one option among many others. Differently from a linear-skip approach, it has a low level of guidance oriented towards linear navigation.</p> </li> </ul> <p><em>Navigation input </em>columns define the ways users can move through the narrative:</p> <ul> <li> <p><code>button_input</code> (boolean)</p> </li> <li> <p><code>scroll_input</code> (boolean)</p> </li> <li> <p><code>slider_input</code> (boolean)</p> </li> </ul> <p><em>Navigation progress </em>columns describe methods through which the reader perceives its placement within the narrative:</p> <ul> <li> <p><code>text_progression</code> (boolean): text or numbers act as signifiers for user position.</p> </li> <li> <p><code>dots_progression</code> (boolean)</p> </li> <li> <p><code>visualisation_progression</code> (boolean): the visualisation used in the narrative, or a visualised progress widget acts as a signifier for user position.</p> </li> </ul> <p><em>Level of control </em>columns describe how much control a reader has over the text, visualisations, and animated transitions. Control could be discrete (D) when it triggers the motion, continuous (C) when it can act throughout all the keyframes, or hybrid (H) if it supports aspects of both. When animation is absent, control can be not available (NA). In particular, while visualisation control is related to the visualisation as a whole (e.g., the entire scatter plot moving up or down the page), the animated transition is related to more specific, data-relevant motion.<br>Columns:</p> <ul> <li> <p><code>text_control</code> (categorical):</p> </li> <ul> <li> <p>D</p> </li> <li> <p>C</p> </li> <li> <p>H</p> </li> </ul> <li> <p><code>visualisation_control</code> (categorical):</p> </li> <ul> <li> <p>D</p> </li> <li> <p>C</p> </li> <li> <p>H</p> </li> </ul> <li> <p><code>animation_control</code> (categorical):</p> </li> <ul> <li> <p>D</p> </li> <li> <p>C</p> </li> <li> <p>H</p> </li> <li> <p>NA</p> </li> </ul> </ul> <p>&nbsp;</p> <h2>References</h2> <p>[1] E. Segel and J. Heer, &ldquo;Narrative Visualization: Telling Stories with Data,&rdquo; IEEE Trans. Visual. Comput. Graphics, vol. 16, no. 6, pp. 1139&ndash;1148, 2010, doi: 10.1109/TVCG.2010.179.</p> <p>[2] M. Terras, J. Nyhan, and E. Vanhoutte, Defining Digital Humanities: A Reader. Routledge, 2016.</p> <p>[3] S. Schreibman, R. G. Siemens, and J. Unsworth, Eds., A companion to digital humanities. in Blackwell companions to literature and culture, no. 26. Malden, MA: Blackwell Pub, 2004.</p> <p>[4] C. Kinkeldey, A. M. MacEachren, and J. Schiewe, &ldquo;How to Assess Visual Communication of Uncertainty? A Systematic Review of Geospatial Uncertainty Visualisation User Studies,&rdquo; The Cartographic Journal, vol. 51, no. 4, pp. 372&ndash;386, 2014, doi: 10.1179/1743277414Y.0000000099.</p> <p>[5] G. Panagiotidou, H. Lamqaddam, J. Poblome, K. Brosens, K. Verbert, and A. Vande Moere, &ldquo;Communicating Uncertainty in Digital Humanities Visualization Research,&rdquo; IEEE Transactions on Visualization and Computer Graphics, vol. 29, no. 1, pp. 635&ndash;645, Jan. 2023, doi: 10.1109/TVCG.2022.3209436.</p> <p>[6] J. Drucker, &ldquo;Humanities Approaches to Graphical Display,&rdquo; Digital Humanities Quarterly, vol. 5, no. 1, 2011, Accessed: Sep. 17, 2024. [Online]. Available: <a href="https://www.digitalhumanities.org/dhq/vol/5/1/000091/000091.html">https://www.digitalhumanities.org/dhq/vol/5/1/000091/000091.html</a></p> <p>[7] F. Windhager et al., &ldquo;Visualization of Cultural Heritage Collection Data: State of the Art and Future Challenges,&rdquo; IEEE Trans. Visual. Comput. Graphics, vol. 25, no. 6, pp. 2311&ndash;2330, Jun. 2019, doi: 10.1109/TVCG.2018.2830759.</p> <p>[8] J. S. Yi, Y. A. Kang, J. Stasko, and J. A. Jacko, &ldquo;Toward a Deeper Understanding of the Role of Interaction in Information Visualization,&rdquo; IEEE Trans. Visual. Comput. Graphics, vol. 13, no. 6, pp. 1224&ndash;1231, 2007, doi: 10.1109/TVCG.2007.70515.</p> <p>[9] E. Dimara and C. Perin, &ldquo;What is Interaction for Data Visualization?,&rdquo; IEEE Transactions on Visualization and Computer Graphics, vol. 26, no. 1, pp. 119&ndash;129, Jan. 2020, doi: 10.1109/TVCG.2019.2934283.</p> <p>[10] M. Brehmer and T. Munzner, &ldquo;A Multi-Level Typology of Abstract Visualization Tasks,&rdquo; IEEE Trans. Visual. Comput. Graphics, vol. 19, no. 12, pp. 2376&ndash;2385, 2013, doi: 10.1109/TVCG.2013.124.</p> <p>[11] Y. Shi, T. Gao, X. Jiao, and N. Cao, &ldquo;Breaking the Fourth Wall of Data Stories Through Interaction,&rdquo; IEEE Trans. Visual. Comput. Graphics, pp. 1&ndash;11, 2022, doi: 10.1109/TVCG.2022.3209409.</p> <p>[12] S. McKenna, N. Henry Riche, B. Lee, J. Boy, and M. Meyer, &ldquo;Visual Narrative Flow: Exploring Factors Shaping Data Visualization Story Reading Experiences,&rdquo; Computer Graphics Forum, vol. 36, no. 3, pp. 377&ndash;387, 2017, doi: 10.1111/cgf.13195.</p> <p>&nbsp;</p> <h2>Fundings</h2> <p>Project funded by the European Union &ndash; NextGenerationEU under the National Recovery and Resilience Plan (NRRP), Investment I.4.1 - Borse PNRR Patrimonio Culturale.</p>

opencc-by-4.0Nov 2024View details →
zenodo52/100

Gravity Spy Machine Learning Classifications of LIGO Glitches from Observing Runs O1, O2, O3a, and O3b

<p>This data set contains all classifications that the Gravity Spy Machine Learning model for LIGO glitches from the first three observing runs (<a href="https://doi.org/10.7935/K57P8W9D">O1</a>, <a href="https://doi.org/10.7935/CA75-FM95">O2</a> and O3, where O3 is split into <a href="https://doi.org/10.7935/nfnt-hm34">O3a</a> and <a href="https://doi.org/10.7935/pr1e-j706">O3b</a>). Gravity Spy classified all noise events identified by the <a href="https://doi.org/10.1016/j.softx.2020.100620">Omicron trigger pipeline</a> in which Omicron identified that the signal-to-noise ratio was above 7.5 and the peak frequency of the noise event was between 10 Hz and 2048 Hz. To classify noise events, Gravity Spy made <a href="https://en.wikipedia.org/wiki/Constant-Q_transform">Omega scans</a> of every glitch consisting of 4 different durations, which helps capture the morphology of noise events that are both short and long in duration.</p> <p>There are <a href="https://doi.org/10.1088/1361-6382/aa5cea">22 classes</a> used for O1 and O2 data (including No_Glitch and None_of_the_Above), while there are <a href="https://doi.org/10.1088/1361-6382/ac1ccb">two additional classes</a> used to classify O3 data (while None_of_the_Above was removed).</p> <p>For O1 and O2, the glitch classes were: 1080Lines, 1400Ripples, Air_Compressor, Blip, Chirp, Extremely_Loud, Helix, Koi_Fish, Light_Modulation, Low_Frequency_Burst, Low_Frequency_Lines, No_Glitch, None_of_the_Above, Paired_Doves, Power_Line, Repeating_Blips, Scattered_Light, Scratchy, Tomte, Violin_Mode, Wandering_Line, Whistle</p> <p>For O3, the glitch classes were: 1080Lines, 1400Ripples, Air_Compressor, Blip, <strong>Blip_Low_Frequency</strong>, Chirp, Extremely_Loud, <strong>Fast_Scattering</strong>, Helix, Koi_Fish, Light_Modulation, Low_Frequency_Burst, Low_Frequency_Lines, No_Glitch, None_of_the_Above, Paired_Doves, Power_Line, Repeating_Blips, Scattered_Light, Scratchy, Tomte, Violin_Mode, Wandering_Line, Whistle</p> <p>The data set is described in <a href="https://doi.org/10.1088/1361-6382/acb633"><strong>Glanzer </strong><em>et al</em><strong>. (2023)</strong></a>, which we ask to be cited in any publications using this data release. Example code using the data can be found in this <a href="https://colab.research.google.com/drive/19q_lItODPk7qw_sohlHyWPnAbY0FZyt8?usp=sharing"><strong>Colab notebook</strong></a>.</p> <p>If you would like to download the Omega scans associated with each glitch, then you can use the gravitational-wave data-analysis tool <a href="https://gwpy.github.io/docs/stable/">GWpy</a>. If you would like to use this tool, please install anaconda if you have not already and create a virtual environment using the following command</p> <pre><code class="language-bash">conda create --name gravityspy-py38 -c conda-forge python=3.8 gwpy pandas psycopg2 sqlalchemy</code></pre> <p>After downloading one of the CSV files for a specific era and interferometer, please run the following Python script if you would like to download the data associated with the metadata in the CSV file. We recommend not trying to download too many images at one time. For example, the script below will read data on Hanford glitches from O2 that were classified by Gravity Spy and filter for only glitches that were labelled as Blips with 90% confidence or higher, and then download the first 4 rows of the filtered table.</p> <pre><code class="language-python">from gwpy.table import GravitySpyTable H1_O2 = GravitySpyTable.read('H1_O2.csv') H1_O2[(H1_O2["ml_label"] == "Blip") &amp; (H1_O2["ml_confidence"] &gt; 0.9)] H1_O2[0:4].download(nproc=1)</code></pre> <p>Each of the columns in the CSV files are taken from various different inputs:&nbsp;</p> <p>[&lsquo;event_time&rsquo;, &lsquo;ifo&rsquo;, &lsquo;peak_time&rsquo;, &lsquo;peak_time_ns&rsquo;, &lsquo;start_time&rsquo;, &lsquo;start_time_ns&rsquo;, &lsquo;duration&rsquo;, &lsquo;peak_frequency&rsquo;, &lsquo;central_freq&rsquo;, &lsquo;bandwidth&rsquo;, &lsquo;channel&rsquo;, &lsquo;amplitude&rsquo;, &lsquo;snr&rsquo;, &lsquo;q_value&rsquo;] contain metadata about the signal from the <a href="https://virgo.docs.ligo.org/virgoapp/Omicron/">Omicron pipeline</a>.&nbsp;</p> <p>[&lsquo;gravityspy_id&rsquo;] is the unique identifier for each glitch in the dataset.&nbsp;</p> <p>[&lsquo;1400Ripples&rsquo;, &lsquo;1080Lines&rsquo;, &lsquo;Air_Compressor&rsquo;, &lsquo;Blip&rsquo;, &lsquo;Chirp&rsquo;, &lsquo;Extremely_Loud&rsquo;, &lsquo;Helix&rsquo;, &lsquo;Koi_Fish&rsquo;, &lsquo;Light_Modulation&rsquo;, &lsquo;Low_Frequency_Burst&rsquo;, &lsquo;Low_Frequency_Lines&rsquo;, &lsquo;No_Glitch&rsquo;, &lsquo;None_of_the_Above&rsquo;, &lsquo;Paired_Doves&rsquo;, &lsquo;Power_Line&rsquo;, &lsquo;Repeating_Blips&rsquo;, &lsquo;Scattered_Light&rsquo;, &lsquo;Scratchy&rsquo;, &lsquo;Tomte&rsquo;, &lsquo;Violin_Mode&rsquo;, &lsquo;Wandering_Line&rsquo;, &lsquo;Whistle&rsquo;] contain the machine learning confidence for a glitch being in a particular Gravity Spy class (the confidence in all these columns should sum to unity). These use the original 22 classes in all cases.</p> <p>[&lsquo;ml_label&rsquo;, &lsquo;ml_confidence&rsquo;] provide the machine-learning predicted label for each glitch, and the machine learning confidence in its classification.&nbsp;</p> <p>[&lsquo;url1&rsquo;, &lsquo;url2&rsquo;, &lsquo;url3&rsquo;, &lsquo;url4&rsquo;] are the links to the publicly-available <a href="https://gwdetchar.readthedocs.io/en/stable/omega/">Omega scans</a> for each glitch. &lsquo;url1&rsquo; shows the glitch for a duration of 0.5 seconds, &lsquo;url2&rsquo; for 1 seconds, &lsquo;url3&rsquo; for 2 seconds, and &lsquo;url4&rsquo; for 4 seconds.</p> <p>For the most recently uploaded training set used in Gravity Spy machine learning algorithms, please see <a href="https://zenodo.org/record/1486046#.YZfcar3MJqs">Gravity Spy Training Set</a> on Zenodo.&nbsp;</p> <p><br> For detailed information on the training set used for the original Gravity Spy machine learning paper, please see <a href="https://zenodo.org/record/1476156#.YZfchL3MJqs">Machine learning for Gravity Spy: Glitch classification and dataset</a> on Zenodo.</p>

opencc-by-4.0Nov 2021View details →
zenodo52/100

Zero Modes and Classification of Combinatorial Metamaterials

<p>This dataset contains the simulation&nbsp;data of the combinatorial metamaterial as used for the paper &#39;Machine Learning of Implicit Combinatorial Rules in Mechanical Metamaterials&#39;, as published in Physical Review Letters.</p> <p>In this paper, the data is used to classify each&nbsp;<span class="math-tex">\(k \times k\)</span> unit cell design into one of two classes (C or I) based on the scaling (linear or constant) of the number of zero modes&nbsp;<span class="math-tex">\(M_k(n)\)</span>&nbsp;for metamaterials consisting of an&nbsp;<span class="math-tex">\(n\times n\)</span>&nbsp;tiling&nbsp;of the corresponding unit cell. Additionally, a random walk&nbsp;through the design space starting from&nbsp;class C unit cells was performed to characterize the boundary between class C and I in design space. A more detailed description of the contents of the dataset follows below.</p> <p><strong>Modescaling_raw_data.zip</strong></p> <p>This file contains uniformly sampled unit cell designs for metamaterial M2&nbsp;and&nbsp;<span class="math-tex">\(M_k(n)\)</span>&nbsp;for&nbsp;<span class="math-tex">\(1\leq n\leq 4\)</span>, which was used to classify the unit cell designs for the data set. There is a small subset of designs for&nbsp;<span class="math-tex">\(k=\{3, 4, 5\}\)</span>&nbsp;that do not neatly fall into the class C and I classification, and instead require additional simulation for&nbsp;<span class="math-tex">\(4 \leq n \leq 6\)</span>&nbsp;before either saturating to a constant number of zero modes (class I) or linearly increasing (class C). This file contains the simulation data of size&nbsp;<span class="math-tex">\(3 \leq k \leq 8\)</span>&nbsp;unit cells. The data is organized as follows.</p> <p>Simulation data for&nbsp;<span class="math-tex">\(3 \leq k \leq 5\)</span>&nbsp;and&nbsp;<span class="math-tex">\(1 \leq n \leq 4\)</span>&nbsp;is stored in numpy array format (.npy) and can be readily loaded in Python with the Numpy package&nbsp;using the numpy.load command. These files are named &quot;data_new_rrQR_i_n_M_kxk_fixn4.npy&quot;, and contain a [Nsim, 1+k*k+4] sized array, where Nsim is the number of simulated unit cells. Each row corresponds to a unit cell. The columns are&nbsp;organized as follows:</p> <ul> <li>col 0: label number to keep track</li> <li>col 1 - k*k+1: flattened unit cell design, numpy.reshape should bring it back to its original&nbsp;<span class="math-tex">\(k \times k\)</span>&nbsp;form.&nbsp;</li> <li>col k*k+1 -&nbsp;k*k+5: number of zero modes&nbsp;<span class="math-tex">\(M_k(n)\)</span>&nbsp;in ascending order of&nbsp;<span class="math-tex">\(n\)</span>, so:&nbsp;<span class="math-tex">\(\{M_k(1), M_k(2), M_k(3), M_k(4)\}\)</span>.</li> </ul> <p><strong>Note:</strong> the unit cell design uses the numbers&nbsp;<span class="math-tex">\(\{0, 1, 2, 3\}\)</span>&nbsp;to refer to each building block orientation. The building block orientations can be characterized through the orientation of the missing diagonal bar (see Fig. 2 in the paper), which can be Left Up (LU), Left Down (LD), Right Up (RU), or Right Down (RD). The numbers correspond to the building block orientation&nbsp;<span class="math-tex">\(\{0, 1, 2, 3\} = \{\mathrm{LU, RU, RD, LD}\}\)</span>.</p> <p>Simulation data for&nbsp;<span class="math-tex">\(3 \leq k \leq 5\)</span>&nbsp;and&nbsp;<span class="math-tex">\(1 \leq n \leq 6\)</span>&nbsp;for unit cells that cannot be classified as class C or I for <span class="math-tex">\(1 \leq n \leq 4\)</span>&nbsp;is stored in numpy array format (.npy) and can be readily loaded in Python with the Numpy package&nbsp;using the numpy.load command. These files are named &quot;data_new_rrQR_i_n_M_kxk_fixn4_classX_extend.npy&quot;, and contain a [Nsim, 1+k*k+6] sized array, where Nsim is the number of simulated unit cells. Each row corresponds to a unit cell. The columns are&nbsp;organized as follows:</p> <ul> <li>col 0: label number to keep track</li> <li>col 1 - k*k+1: flattened unit cell design, numpy.reshape should bring it back to its original&nbsp;<span class="math-tex">\(k \times k\)</span>&nbsp;form.&nbsp;</li> <li>col k*k+1 -&nbsp;k*k+5: number of zero modes&nbsp;<span class="math-tex">\(M_k(n)\)</span>&nbsp;in ascending order of&nbsp;<span class="math-tex">\(n\)</span>, so:&nbsp;<span class="math-tex">\(\{M_k(1), M_k(2), M_k(3), M_k(4), M_k(5), M_k(6)\}\)</span>.</li> </ul> <p>Simulation data for&nbsp;<span class="math-tex">\(6 \leq k \leq 8\)</span>&nbsp;&nbsp;unit cells are&nbsp;stored in numpy array format (.npy) and can be readily loaded in Python with the Numpy package&nbsp;using the numpy.load command. Note that the number of modes is now calculated for&nbsp;<span class="math-tex">\(n_x \times n_y\)</span>&nbsp;metamaterials, where we calculate&nbsp;<span class="math-tex">\((n_x, n_y) = \{(1,1), (2, 2), (3, 2), (4,2), (2, 3), (2, 4)\}\)</span>&nbsp;rather than&nbsp;<span class="math-tex">\(n_x=n_y=n\)</span>&nbsp;to save computation time.&nbsp;These files are named &quot;data_new_rrQR_i_n_Mx_My_n4_kxk(_extended).npy&quot;, and contain a [Nsim, 1+k*k+8] sized array, where Nsim is the number of simulated unit cells. Each row corresponds to a unit cell. The columns are&nbsp;organized as follows:</p> <ul> <li>col 0: label number to keep track</li> <li>col 1 - k*k+1: flattened unit cell design, numpy.reshape should bring it back to its original&nbsp;<span class="math-tex">\(k \times k\)</span>&nbsp;form.&nbsp;</li> <li>col k*k+1 -&nbsp;k*k+9: number of zero modes&nbsp;<span class="math-tex">\(M_k(n_x, n_y)\)</span>&nbsp;in order:&nbsp;<span class="math-tex">\(\{M_k(1, 1), M_k(2, 2), M_k(3, 2), M_k(4, 2), M_k(1, 1), M_k(2, 2), M_k(2, 3), M_k(2, 4)\}\)</span>.</li> </ul> <p>Simulation data of metamaterial M1 for <span class="math-tex">\(k_x \times k_y\)</span> metamaterials are stored in compressed numpy array format (.npz) and can be loaded in Python with the Numpy package using the numpy.load command. These files are named &quot;smiley_cube_x_y_<span class="math-tex">\(k_x\)</span>x<span class="math-tex">\(k_y\)</span>.npz&quot;, which contain all possible metamaterial designs, and &quot;smiley_cube_uniform_sample_x_y_<span class="math-tex">\(k_x\)</span>x<span class="math-tex">\(k_y\)</span>.npz&quot;, which contain uniformly sampled metamaterial designs. The configurations are accessed with the keyword argument &#39;configs&#39;. The classification is accessed with the keyword argument &#39;compatible&#39;. The configurations array is of shape [Nsim, <span class="math-tex">\(k_x\)</span>, <span class="math-tex">\(k_y\)</span>], the classification array is of shape [Nsim]. The building blocks in the configuration are denoted by 0 or 1, which correspond to the red/green and white/dashed building blocks respectively. Classification is 0 or 1, which corresponds to I and C respectively.</p> <p><strong>Modescaling_classification_results.zip</strong></p> <p>This file contains the classification, slope, and offset of the scaling of the number of zero modes&nbsp;<span class="math-tex">\(M_k(n)\)</span>&nbsp;for the unit cells of metamaterial M2 in&nbsp;Modescaling_raw_data.zip. The data is organized as follows.</p> <p>The results for&nbsp;<span class="math-tex">\(3 \leq k \leq 5\)</span>&nbsp;based on the&nbsp;<span class="math-tex">\(1 \leq n \leq 4\)</span>&nbsp;mode scaling data is stored in &quot;results_analysis_new_rrQR_i_Scen_slope_offset_M1k_kxk_fixn4.txt&quot;. The data can be loaded using &#39;,&#39; as delimiter. Every row corresponds to a unit cell design (see the label number to compare to the earlier data). The columns are organized as follows:</p> <p>col 0: label number to keep track</p> <p>col 1: the class, where 0 corresponds to class I, 1 to class C and 2 to class X (neither class I or C for&nbsp;<span class="math-tex">\(1 \leq n \leq 4\)</span>)</p> <p>col 2: slope from&nbsp;<span class="math-tex">\(n \geq 2\)</span>&nbsp;onward (undefined for class X)</p> <p>col 3: the offset is defined as&nbsp;<span class="math-tex">\(M_k(2) - 2 \cdot \mathrm{slope}\)</span></p> <p>col 4:&nbsp;<span class="math-tex">\(M_k(1)\)</span></p> <p>The results for&nbsp;<span class="math-tex">\(3 \leq k \leq 5\)</span>&nbsp;based on the extended&nbsp;<span class="math-tex">\(1 \leq n \leq 6\)</span>&nbsp;mode scaling data is stored in &quot;results_analysis_new_rrQR_i_Scen_slope_offset_M1k_kxk_fixn4_classC_extend.txt&quot;. The data can be loaded using &#39;,&#39; as delimiter. Every row corresponds to a unit cell design (see the label number to compare to the earlier data). The columns are organized as follows:</p> <p>col 0: label number to keep track</p> <p>col 1: the class, where 0 corresponds to class I, 1 to class C and 2 to class X (neither class I or C for <span class="math-tex">\(1 \leq n \leq 6\)</span>)</p> <p>col 2: slope from&nbsp;<span class="math-tex">\(n \geq 2\)</span>&nbsp;onward (undefined for class X)</p> <p>col 3: the offset is defined as&nbsp;<span class="math-tex">\(M_k(2) - 2 \cdot \mathrm{slope}\)</span></p> <p>col 4:&nbsp;<span class="math-tex">\(M_k(1)\)</span></p> <p>The results for&nbsp;<span class="math-tex">\(6 \leq k \leq 8\)</span>&nbsp;based on the&nbsp;<span class="math-tex">\(1 \leq n \leq 4\)</span>&nbsp;mode scaling data is stored in &quot;results_analysis_new_rrQR_i_Scenx_Sceny_slopex_slopey_offsetx_offsety_M1k_kxk(_extended).txt&quot;. The data can be loaded using &#39;,&#39; as delimiter. Every row corresponds to a unit cell design (see the label number to compare to the earlier data). The columns are organized as follows:</p> <p>col 0: label number to keep track</p> <p>col 1: the class_x based on <span class="math-tex">\(M_k(n_x, 2)\)</span>, where 0 corresponds to class I, 1 to class C and 2 to class X (neither class I or C for <span class="math-tex">\(1 \leq n_x \leq 4\)</span>)</p> <p>col 2: the class_y based on <span class="math-tex">\(M_k(2, n_y)\)</span>, where 0 corresponds to class I, 1 to class C and 2 to class X (neither class I or C for <span class="math-tex">\(1 \leq n_y \leq 4\)</span>)</p> <p>col 3: slope_x from&nbsp;<span class="math-tex">\(n_x \geq 2\)</span>&nbsp;onward (undefined for class X)</p> <p>col 4: slope_y from&nbsp;<span class="math-tex">\(n_y \geq 2\)</span>&nbsp;onward (undefined for class X)</p> <p>col 5: the offset_x is defined as&nbsp;<span class="math-tex">\(M_k(2, 2) - 2 \cdot \mathrm{slope_x}\)</span></p> <p>col 6: the offset_x is defined as&nbsp;<span class="math-tex">\(M_k(2, 2) - 2 \cdot \mathrm{slope_y}\)</span></p> <p>col 7:&nbsp;<span class="math-tex">\(M_k(1, 1)\)</span></p> <p>Additionally, results including classification for M2.ii can be found in the &quot;results_analysis_unimodal_vs_oligomodal_vs_plurimodal_i_Scen_slope_M_M1k_kxk.txt and &quot;results_analysis_unimodal_vs_oligomodal_vs_plurimodal_i_Scenx_Sceny_slopex_slopey_Mx_My_M1k_kxk.txt&quot; files.</p> <p><strong>Random Walks Data</strong></p> <p>This file contains the random walks for&nbsp;<span class="math-tex">\(3 \leq k \leq 8\)</span>&nbsp;unit cells of metamaterial M2. The random walk starts from a class C unit cell design (classification M2.ii), for each step&nbsp;<span class="math-tex">\(s\)</span>&nbsp;a randomly picked unit cell is changed to a random new orientation for a total of&nbsp;<span class="math-tex">\(s=k^2\)</span>&nbsp;steps. The data is organized as follows.</p> <p>The configurations for each step are stored in the files named &quot;configlist_test_i.npy&quot;, where i is a number and corresponds to a different starting unit cell. The stored array has the shape [k*k+1, 2*k+2, 2*k+2]. The first dimension denotes the step&nbsp;<span class="math-tex">\(s\)</span>, where&nbsp;<span class="math-tex">\(s=0\)</span>&nbsp;is the initial configuration. The second and third dimension denote the unit cell configuration in the pixel representation (see paper) padded with a single pixel wide layer using periodic boundary conditions.&nbsp;</p> <p>The class for each configuration are stored in &quot;lmlist_test_i.npy&quot;, where i corresponds to the same number as for the configurations in the &quot;configlist_test_i.npy&quot; file. The stored&nbsp;array has the shape [k*k+1], where the index corresponds to the step&nbsp;<span class="math-tex">\(s\)</span>&nbsp;and displays the class for the accompanying unit cell. The stored number corresponds to the class as&nbsp;<span class="math-tex">\(\{0, 1\} = \{\mathrm{I}, \mathrm{C}\}\)</span>.&nbsp;</p>

opencc-by-4.0Feb 2022View details →
zenodo52/100

Meta-analysis and gender classification of 914 national and international surveys in six European countries (2000-2023)

<p><span>This data frame presents the results of a quan</span><span>ti</span><span>ta</span><span>ti</span><span>ve content analysis of the occurrence of gender‐based concepts, themes, issues, and solu</span><span>ti</span><span>ons within large‐scale poli</span><span>ti</span><span>cal and sociological survey ques</span><span>ti</span><span>onnaires fielded cross‐na</span><span>ti</span><span>onally in Europe and in six European countries: Denmark, Germany, Hungary, Switzerland and the UK, spanning 2000‐2023. Data was collected by teams from each country between September 2023‐January 2024. Teams collected ques</span><span>ti</span><span>ons in the original language and provided a transla</span><span>ti</span><span>on into English. Analysis was conducted using the translated text. The unit of analysis (&lsquo;CODING_UNIT_TEXT&rsquo;) was the individual 'gender‐related argument' within a survey ques</span><span>ti</span><span>on. This could be the en</span><span>ti</span><span>re survey ques</span><span>ti</span><span>on, a sub‐ques</span><span>ti</span><span>on (in the case of matrix ques</span><span>ti</span><span>ons), or a singular response op</span><span>ti</span><span>on (for mul</span><span>ti</span><span>ple choice ques</span><span>ti</span><span>ons). Coding units were coded in three key domains:(1) Gender concepts, (2) Themes/issues, and (3) Solu</span><span>ti</span><span>ons. Up to two Themes/Issues and Solu</span><span>ti</span><span>ons could be coded per coding unit. Several coding categories within the Themes/Issues and Solu</span><span>ti</span><span>ons domains func</span><span>ti</span><span>on hierarchically, where a coder first assigned a higher‐level category and then as many subcategories as applicable. For example, a ques</span><span>ti</span><span>on concerning government‐funded childcare is coded as B1_Economy ‐&gt; B1_4_LabourMarket ‐&gt; B1_4_1_CareWork ‐&gt; B1_4_1_3_Childcare. The corresponding codebook presents the uni</span><span>ti</span><span>sa</span><span>ti</span><span>on process and coding categories in full detail.</span></p>

opencc-by-sa-4.0Jun 2024View details →
zenodo52/100

Holdridge Life Zones Classification in New Caledonia Habitats

<h1>Description</h1> <p>This dataset aims to represent, in geographic space, the distribution of life zones as first defined by Holdridge in 1947 and updated in 1967. Life zones are delineated through three parameters:</p> <ul> <li>Mean Annual Biotemperature (&deg;C): This axis represents the average annual temperature, considering only temperatures above 0&deg;C, as it influences biological activity. It determines the thermal regime of the environment.</li> <li>Annual Precipitation (mm): This axis measures the total annual precipitation, indicating moisture availability. It is crucial for determining the hydric regime and supporting different types of vegetation and ecosystems.</li> <li>Potential Evapotranspiration Ratio (PET): This axis is the ratio of potential evapotranspiration to annual precipitation. It reflects the balance between water demand and supply, indicating aridity or humidity levels and influencing vegetation types and ecosystem dynamics.</li> </ul> <p>We used a combination of WorldClim datasets (Biotemperature and potential evapotranspiration) and M&eacute;t&eacute;o-France Aurelhy datasets (Annual Precipitation) specifically designed for New Caledonia to produce the raster with a 1 km&sup2; resolution.</p> <h1>Content</h1> <p>This dataset was produced, analyzed, and verified using a combination of open-source software, including QGIS, PostgreSQL, PostGIS, Python, R and the GDAL library, all running on Linux.</p> <ul> <li>amap_raster_holdridge_nc.tif is a GeoTIFF, utilizing the WGS84 international coordinate system, and consists of a single band with three major classes coded as <ul> <li>Dry life zone (rast = 1)</li> <li>Moist life zone (rast = 2)</li> <li>Rain life zone (rast = 3)</li> </ul> </li> <li>holdridge_3classes_NC.png is an image illustrating the valid domain of life zones in New Caledonia and the classification used in the dataset.</li> </ul> <h1>Limitations</h1> <p>Strictly, the classification leads to five distinct classes (very dry, dry, moist, wet, and rain), but as the two extreme classes cover less than 0.5% of New Caledonia, we merged very dry and dry into the "dry" class, as well as wet and rain into the "rain" class as illustrated in the Figure <a href="../api/records/12731521/draft/files/holdridge_3classes_NC.png/content" target="_blank" rel="noopener noreferrer">holdridge_3classes_NC.png</a>.</p>

opencc-by-4.0Jul 2024View details →
zenodo52/100

WorldSeasons: a seasonal classification system interpolating biomes within the year for improved temporal aggregation

<p>We present a seasonal classification system to improve the temporal framing of comparative scientific analysis. Research often uses yearly aggregates to understand inherently seasonal phenomena like harvests, monsoons, and droughts. This obscures important trends across time and differences through space by including redundant data. Our classification system allows for a more targeted approach. We split global land into four principal climate zones: desert, arctic and high montane, tropical, and temperate. A cluster analysis with zone-specific variables and weighting splits each month of the year into discrete seasons based on the monthly climate. We expect the data will be able to answer global comparative analysis questions like: are global winters less icy than before? Are wildfires more frequent now in the dry season? How severe are monsoon season flooding events? This is a natural extension of the historical concept of biomes, made possible by recent advances in climate data availability and artificial intelligence.</p>

opencc-by-4.0Aug 2024View details →
zenodo52/100

Dataset of "Anomaly Detection in Industrial Networks: Current State, Classification, and Key Challenges"

<p>Industrial networks are adapted to their specific requirements, especially in terms of industrial processes. To ensure sufficient security in these networks, it is necessary to set and use security policies that complement government regulations, recommendations, and relevant security standards. This paper aims to provide an in-depth analysis of the anomalies occurring within the networks and propose a structure for collecting valuable data from the experimental site based on dividing anomalies into three main categories:<br>security, operational, and service anomalies (and regular traffic recognition). We present a proof-of-concept solution/design aggregating data in industrial networks for advanced anomaly classification. Multiple data sources such as industrial communication, sensor data (additional sensors controlling device behavior), and HW status data are used as data sources. A total of three scenarios (using a physical testbed) were implemented, where we achieved an accuracy of 0.8540/0.9972 in advanced anomaly classification.</p>

opencc-by-4.0Aug 2024View details →
zenodo52/100

Crowd4SDG - Crowdsourced image classification and damage assessment

<p>This data set contains crowdsourced classification and damage assessment of images of an earthquake extracted from social media.&nbsp;&nbsp;<br> <br> A data set of 907 images posted on Twitter related to the 2019 Albanian Earthquake,&nbsp;that are filtered and pre-classified using an automated technique is cross-validated for accuracy by two different crowds. One, digital humanitarian volunteers using the crowdsourcing platform <a href="http://www.crowd4ems.org">CROWD4EMS</a>&nbsp;and another, paid micro-taskers of&nbsp;the Amazon Mechanical Turk. In order to compare and evaluate the efficiency and accuracy of the volunteers and the paid micro taskers, ground truth is established with the help of a team of experts, who validated the same set of data.&nbsp;<br> <br> <strong>Parameters considered for volunteer contributions:</strong> The dataset was imported to the Crowd4EMS platform for Crowd contribution. In the forum, each volunteer will see the image to be validated along with the tweet text and the link to the original tweet. The user has to validate whether the given image is <em>relevant or</em>&nbsp;<em>irrelevant</em> to the disaster. In case of doubt, the user can refer to the tutorial explaining the relevance or skip the task. Once the image&#39;s relevance is validated, the user will be asked to label the <em>severity</em> of the impact, as seen in the image.</p> <p>The Automated algorithm has pre-classified the images as <em>severe&nbsp;</em>and <em>minimal </em>damage. The Crowd4EMS platform lets the volunteer label them as &#39;<em>severe damage</em>,&#39;&nbsp;<em>moderate damage&#39;</em>,&#39;&nbsp;<em>minimal damage&#39;,&nbsp;</em>and&#39;&nbsp;<em>no damage&#39;.</em>&nbsp;Each task has to be answered <em>at least three times</em>, and the final consensus is taken as per the<em> inter-rater agreement.&nbsp;</em><br> <br> <strong>Parameters considered for micro-taskers contribution:</strong>&nbsp;The dataset was imported to the <em>Amazon Mechanical Turk</em> platform for Crowd contribution. In the platform, each worker will see only the image that is to be categorised as follows:&nbsp;The user has to validate whether the given image depicts&nbsp;<em>severe damage, moderate damage, minimal damage, no damage&nbsp;</em>or&nbsp;<em>irrelevant</em> to the disaster. Each task has to be answered <em>at least ten times</em>, and the final consensus is taken as per the<em> inter-rater agreement.&nbsp;</em><br> <br> <strong>Acknowledgements:</strong> We want to thank Muhammad Imran&nbsp;of&nbsp;Qatar Computing Research Institute for sharing their pre-filtered social media imagery dataset on the Albanian earthquake from the Artificial Intelligence for Disaster Response (AIDR) Platform.&nbsp;We would also like to extend our gratitude to the volunteers for their contribution on the Crowd4EMS Platform.<br> &nbsp;</p>

opencc-by-4.0Sep 2021View details →
zenodo52/100

EUNIS-ESy: Expert system for automatic classification of European vegetation plots to EUNIS habitats

<p><strong>EUNIS-ESy</strong> is an expert system for automatic classification of European vegetation plots to habitat types of the EUNIS Habitat Classification. The EUNIS classification and the principles of the expert system are described by <a href="https://doi.org/10.1111/avsc.12519">Chytr&yacute; et al. (2020)</a>. The classification of a set of vegetation plots can be run using the&nbsp;JUICE program (<a href="https://doi.org/10.1111/j.1654-1103.2002.tb02069.x">Tich&yacute; 2002</a>; <a href="https://www.sci.muni.cz/botany/juice/">https://www.sci.muni.cz/botany/juice/</a>), TURBOVEG 3 program (Hennekens 2015) and an R script (<a href="https://doi.org/10.1111/avsc.12562">Bruelheide et al. 2021</a>).</p> <p>This dataset contains two parts: (1) the expert system and related files necessary for running it; (2) characterization of EUNIS habitats based on the results of the expert system classification.</p> <p><strong>1. Expert system and related files necessary to run it</strong></p> <p>1.1. <strong>EUNIS-ESy-2025-10-03.txt </strong>&ndash; a file containing the script for the classification of vegetation plots by EUNIS-ESy. This version contains tested definitions for the revised EUNIS classification of vegetated Marine (MA), Coastal (N), Wetland (Q), Grassland (R), Shrubland (S), Forest (T), Inland sparsely vegetated (U) and Man-made (V). It also contains tested definitions of Aquatic plant communities (P3) and Springs (P2N). This file is different from the analogous file in the previous versions.</p> <p>1.2.&nbsp;<strong>Nomenclature-translation-from-Turboveg-2-databases.zip </strong>&ndash; an archive containing the scripts for automatic translation of taxon concepts and names used in individual European Turboveg 2 databases (<a href="https://doi.org/10.2307/3237010">Hennekens &amp; Schamin&eacute;e 2001</a>;&nbsp;<a href="https://www.synbiosys.alterra.nl/turboveg/">https://www.synbiosys.alterra.nl/turboveg/</a>) to the nomenclature that can be used as an input for EUNIS-ESy. This file is the same as in the previous versions.</p> <p>1.3. <strong>EUNIS-ESy-User-Guide.pdf </strong>&ndash; a brief user guide to the classification of vegetation plots by EUNIS-ESy using the JUICE program. Please read this guide carefully before running the expert system to avoid misclassifications. This file is the same as in the previous versions.</p> <p><strong>2. Characterization of the EUNIS habitats based on the results of the EUNIS-ESy classification</strong></p> <p>2.1. <strong>EUNIS-habitats-2025-10-03.xlsx </strong>&ndash; the current list of EUNIS habitats. This file is different from the analogous file in the previous versions.</p> <p>2.2. <strong>EUNIS-EuroVegChecklist-crosswalk-2025-10-03.xlsx</strong> &ndash; a crosswalk between the EUNIS habitat classification and phytosociological alliances of EuroVegChecklist (<a href="http://doi.org/10.1111/avsc.12257">Mucina et al. 2016</a>; <a href="https://floraveg.eu/vegetation/">https://floraveg.eu/vegetation/</a>).</p> <p>2.3.&nbsp;<strong>EUNIS-habitats-Characteristic-species-combintation-2025-10-03.xlsx </strong>&ndash; a database of habitats' characteristic species combinations in a spreadsheet format. These species combinations are based on the analysis of vegetation plots from the European Vegetation Archive (EVA;&nbsp;<a href="https://doi.org/10.1111/avsc.12191">Chytr&yacute; et al. 2016</a>; <a href="http://euroveg.org/eva-database">http://euroveg.org/eva-database</a>) and other databases classified by EUNIS-ESy v2025-10-03. Analytical methods are described in <a href="https://doi.org/10.1111/avsc.12519">Chytr&yacute; et al. (2020)</a>. This file is different from the analogous file in the previous versions.</p> <p>2.4. <strong>EUNIS-habitats-Distribution-maps-2025-10-03.xlsx </strong>&ndash; a set of distribution maps in the TIFF format based on the analysis of vegetation plots from the European Vegetation Archive (EVA;&nbsp;<a href="https://doi.org/10.1111/avsc.12191">Chytr&yacute; et al. 2016</a>; <a href="http://euroveg.org/eva-database">http://euroveg.org/eva-database</a>) and other databases classified by EUNIS-ESy v2025-10-03.</p> <p>2.5.&nbsp;<strong>Data-sources-EUNIS-classification-2025-10-03.pdf </strong>&ndash; a list of data sources used to produce the distribution maps and characteristic species combinations.</p> <p>-----------------------------------------------------------------------------------------------------</p> <p><strong>Differences from the previous version (2021-06-01)</strong></p> <p>Aquatic plant communities (P3), spring (P2N), some wetland (Q61-Q63) and some inland sparsely vegetated (U71-U72) habitats were added to the EUNIS-ESy expert system. Plant taxon concepts and nomenclature were extensively revised. Some previously included habitat definitions were slightly refined. New vegetation-plot records added to the EVA database by 8 August 2025 were used to characterize habitat types. Unlike in the previous version, this version does not provide Habitat factsheets because summarized information about each habitat is now available in the FloraVeg.EU database at <a href="https://floraveg.eu/habitat/">https://floraveg.eu/habitat/</a>.</p> <p>-----------------------------------------------------------------------------------------------------</p> <p>&nbsp;</p> <p><strong>Recommended citation of this version of the EUNIS-ESy expert system</strong></p> <p>Chytr&yacute; et al. (2020), version 2025-10-03</p> <p>Chytr&yacute; M., Tich&yacute; L., Hennekens S.M., Knollov&aacute; I., Janssen J.A.M., Rodwell J.S., Peterka T., Marcen&ograve; C., Landucci F., Danihelka J., H&aacute;jek M., Dengler J., Nov&aacute;k P., Zukal D., Jim&eacute;nez-Alfaro B., Mucina L., Abdulhak S., Aćić S., Agrillo E., Attorre F., Bergmeier E., Biurrun I., Boch S., B&ouml;l&ouml;ni J., Bonari G., Braslavskaya T., Bruelheide H., Campos J.A., Čarni A., Casella L., Ćuk M., Ću&scaron;terevska R., De Bie E., Delbosc P., Demina O., Didukh Y., D&iacute;tě D., Dziuba T., Ewald J., Gavil&aacute;n R.G., G&eacute;gout J.-C., Giusso del Galdo G.P., Golub V., Goncharova N., Goral F., Graf U., Indreica A., Isermann M., Jandt U., Jansen F., Jansen J., Ja&scaron;kov&aacute; A., Jirou&scaron;ek M., Kącki Z., Kaln&iacute;kov&aacute; V., Kavgacı A., Khanina L., Korolyuk A.Yu., Kozhevnikova M., Kuzemko A., K&uuml;zmič F., Kuznetsov O.L., Laiviņ&scaron; M., Lavrinenko I., Lavrinenko O., Lebedeva M., Lososov&aacute; Z., Lysenko T., Maciejewski L., Mardari C., Marin&scaron;ek A., Napreenko M.G., Onyshchenko V., P&eacute;rez-Haase A., Pielech R., Prokhorov V., Ra&scaron;omavičius V., Rodr&iacute;guez Rojo M.P., Rūsiņa S., Schrautzer J., &Scaron;ib&iacute;k J., &Scaron;ilc U., &Scaron;kvorc Ž., Smagin V.A., Stančić Z., Stanisci A., Tikhonova E., Tonteri T., Uogintas D., Valachovič M., Vassilev K., Vynokurov D., Willner W., Yamalov S., Evans D., Palitzsch Lund M., Spyropoulou R., Tryfon E., Schamin&eacute;e J.H.J. (2020) EUNIS Habitat Classification: expert system, characteristic species combinations and distribution maps of European habitats. Applied Vegetation Science, 23, 648&ndash;675. https://doi.org/10.1111/avsc.12519</p>

opencc-by-4.0Dec 2019View details →
zenodo52/100

Global distribution of predicted soil types at 1 km resolution based on the WRB 2022 classification

<p>Global maps at 1 km spatial resolution of the predicted soil types (0&ndash;100% probabilities) at 1 km resolution based on the <a href="https://www.fao.org/soils-portal/data-hub/soil-classification/world-reference-base/en/">WRB 2022</a> (<strong>World Reference Base</strong> the international standard for soil classification) classification system. The training data comes from the following 3 main sources:</p> <ol> <li>WOSIS points available via: <a href="https://www.isric.org/explore/wosis">https://www.isric.org/explore/wosis</a>;</li> <li>HWSD v2 (random draw of cca 20,000 points): <a href="https://iiasa.ac.at/models-tools-data/hwsd">https://iiasa.ac.at/models-tools-data/hwsd</a>;</li> <li>Other national datasets / data from publications and projects.</li> </ol> <p>Predictions are based on using Rando Forest algorithm as implemented in the <a href="https://www.randomforestsrc.org/">randomForestSRC package</a> with cca 190 covariate layers representing soil forming factors (CHELSA Climate, Global Lithological DB GLiM, MODIS EVI and LST long-term derivatives, Digital Terrain model parameters and similar).</p> <p>All TIF files are provided as <a href="https://www.cogeo.org/">COGs</a>, which means that you can open them directly in QGIS or similar.&nbsp;Publication explaining all modeling steps is pending.</p> <p>Update of the predictions takes about 4&ndash;5 hrs and will be regularly run provided that new training points are available. Disclaimer: These are initial results with limited accuracy and possible issues with quality of training points, location errors and harmonization issues. Use at own risk.</p> <p>Note: original list of soil types have been subset to classes that appear at least 10 times and at least in 2 countries. If you notice an error or artifact <strong>please report via <a href="https://github.com/OpenGeoHub/SoilTypeMapping">the Github repository</a></strong>. Help us improve this dataset by contributing training points.</p>

opencc-by-4.0Apr 2023View details →
edi52/100

Landscape Ecosystem Classification Soils and Vegetation Plots Data at the University of Michigan Biological Station, Pellston, Michigan from 1987 to 2015 remeasurements

Landscape ecosystems are a means of understanding the spatial patterns of and the functional interrelationships in forest ecosystems. Landscape ecosystem research is a multifactor, holistic approach to identifying, classifying, describing, and mapping terrain ecosystems. Abiotic and biotic factors are integrated in the field to distinguish repeating units similar in ecological structure and function. Landscape ecosystems are identified by simultaneous integration of physiographic, soil, and vegetation information. The more stable components--physiography and soil--largely determine local climate, and water and nutrient relations, and thus the interrelationships of physiography and soil form the foundation of a landscape ecosystem classification. Vegetation is seen as a phytometer that integrates the many abiotic factors and their interactions, and therefore reflects differences in ecosystem structure and function. When the three main ecosystem factors are analyzed simultaneously, one can perceive interrelationships that result in ecologically meaningful differences among segments of the ecosphere. Landscape ecosystems are spatial; they are volumetric, multi-dimensional segments of earth, whose components include soil, water, atmosphere, solar radiation, and biota. These segments can be identified, classified, described, and mapped at various scales. From the years of 1988 to 2001, various graduate students of Burton V. Barnes completed their masters thesis and dissertations in this pursuit. The attached data set is a culmination of these individual work. Each plot has measurements at various scales within the 10 by 30 meet plot. A stratified random design was used to locate plot locations. The random design was stratified by major and minor landforms in the region. All trees within the plot where identified and dbh was measured. All individual shrubs where identified and abundance was counted within the entire plot. Soils pits locations for each plot where selected

openCC (other)Apr 2025View details →
edi52/100

Land use and land cover (LULC) classification of the CAP LTER study area (central Arizona, USA) using Landsat imagery: 2015 and 2020

## overview The project extends the long-term, LULC datasets to facilitate environmental change monitoring and social-ecological studies regarding urban sprawl and dynamics, urban heat islands, and outdoor water consumption, among others. Six land-use/land-cover (LULC) maps at 30 m resolution were previously created from 1985 to 2010 at five-year intervals (Zhang and Li 2017). This project updates that suite with maps for 2015 and 2020. As with the prior set, systematic object-based classification was utilized to ensure map consistency and direct comparison capability over time. The maps comprise 11 land-use/land-cover classes with an overall accuracy of 89.1% for 2015 and 89.6% for 2020. ## literature cited - Zhang, Y. and X. Li. 2017. Land cover classification of the CAP LTER study area at five-year intervals from 1985 to 2010 using Landsat imagery ver 1. Environmental Data Initiative. https://doi.org/10.6073/pasta/dab4db27974f6c8d5b91a91d30c7781d (Accessed 2022-07-13).

openCC0Aug 2023View details →
zenodo48/100

CzechVeg-ESy: Expert system for automatic classification of vegetation plots from the Czech Republic

<p><strong>Expertn&iacute; syst&eacute;m pro automatickou klasifikaci fytocenologick&yacute;ch sn&iacute;mků z Česk&eacute; republiky&nbsp;</strong><br> [popis a instrukce v če&scaron;tině jsou uvedeny n&iacute;že]</p> <p>***************************************************************************************************************</p> <p><strong>CzechVeg-ESy</strong> is an expert system for automatic classification of vegetation plots from the Czech Republic to the vegetation types defined in the monograph <em>Vegetation of the Czech Republic</em> (<a href="https://www.sci.muni.cz/botany/vegsci/vegetace.php?page=monograph&amp;lang=en">Chytr&yacute; 2007-2013</a>). It is delivered in two main versions: The <strong>main version 1 (v1) </strong>is the original version used for the vegetation classification in that monograph, described in detail in its English Summary (<a href="https://www.sci.muni.cz/botany/chytry/Vegetation-Czech-Rep-Summary.pdf">Chytr&yacute; 2007</a>). Its subversions (indicated by dates) contains corrections of minor errors and species nomenclature. This version only classifies vegetation to the phytosociological associations. The <strong>main version 2 (v2)</strong>&nbsp;uses the same classification system as accepted in the&nbsp;<em>Vegetation of the Czech Republic</em>, but includes&nbsp;more advanced functions to provide more accurate classification. Moreover, it is hierarchical, performing classification not only to associations but also to alliances and classes (for the plots not classified at lower levels).</p> <p>Each version is delivered in two variants. The <strong>basic variant </strong>(CzechVeg-ESy-basic-v1.txt)&nbsp;assigns vegetation plots to associations following their formal definitions created using the Cocktail method (<a href="https://doi.org/10.2307/3236796">Bruelheide 2000</a>) modified by <a href="https://www.sci.muni.cz/botany/chytry/Koci_etal2003_JVS.pdf">Koč&iacute; et al (2003)</a>.&nbsp;These definitions&nbsp;are based on the presence of sociological species groups and the dominance of selected species. The expert system evaluates individual vegetation plots and assigns them to the associations. A classification process is considerably faster when the basic variant is used, as opposed to the full variant. The <strong>full variant</strong> (file&nbsp;CzechVeg-ESy-full-v1.txt) performs the same functions as the basic variant, but in addition, it can also assign the plots not classified by formal definitions based on their numerical similarity to the plots that fulfil&nbsp;the requirements of the formal definitions. Most of such plots can be considered as untypical from the phytosociological point of view, i.e. with poor correspondence to any of the defined vegetation types, in most cases because of the lack of ecologically specialized species. The method of similarity-based assignment is the Frequency-Positive Fidelity Index (FPFI) described in <a href="https://www.sci.muni.cz/botany/chytry/Koci_etal2003_JVS.pdf">Koč&iacute; et al. (2003)</a> and <a href="https://doi.org/10.1007/s11258-004-5798-8">Tich&yacute; (2005)</a>.</p> <p><strong>Instructions for using CzechVeg-ESy version 1</strong></p> <ol> <li>Vegetation plots have to be stored in the TURBOVEG 2 program (<a href="https://www.synbiosys.alterra.nl/turboveg/">https://www.synbiosys.alterra.nl/turboveg/</a>) with the species list Czechia-Slovakia-2015 (contained in the file&nbsp;TurbovegSlBackup_Czechia_slovakia_2015.zip), which is largely compatible with the older species lists called Czechia-Slovakia-2012 and Central-Europe.</li> <li>Export plots from&nbsp;TURBOVEG 2 to a CC! file (Export / Other formats / JUICE input files) and import this file to the JUICE program (<a href="https://www.sci.muni.cz/botany/juice/">https://www.sci.muni.cz/botany/juice/</a>) using the species list in the file&nbsp;Checklist-Danihelka-et-al-2012-ver-2019-07-06.txt, which converts the plant nomenclature to correspond with the Checklist of vascular plants of the Czech Republic (<a href="http://www.preslia.cz/P123Danihelka.pdf">Danihelka et al. 2012</a>).</li> <li>Select Analysis / Expert system in the JUICE program.</li> <li>Upload the expert system file (either&nbsp;CzechVeg-ESy-basic-v1.txt or&nbsp;CzechVeg-ESy-full-v1.txt)&nbsp;by pressing Load ES File button.</li> <li>Modify species nomenclature by pressing the Modify Species Names button.</li> <li>If there are juvenile species in the herb layer in the plots to be analysed, delete them&nbsp;by pressing Delete Juveniles button.</li> <li>In some cases, narrow species concepts were changed to broader concepts, resulting in repetitions of the same names in the table. These must be merged using the Merge Same Spec. Names button. This will also merge records of the same species in different vegetation layers because the expert system assumes that each species name is contained only once in the same plot.</li> <li>When using the full version of the expert system, a threshold similarity value for similarity-based assignment must be specified in the field in the bottom right part of the form. The higher the value, the fewer plots will be assigned, while only those plots will be assigned that have a high similarity to the association. If the threshold value is set to 0, all the plots will be assigned, but some of them may be very dissimilar to the associations which they are assigned to.</li> <li>Run expert system using the Classify Relev&eacute; button (a plot marked by a previous mouse click will be classified) or Classify [colour] Relev&eacute;s button (all the plots of the selected colour will be assigned).</li> <li>If a single plot is classified, species groups present in this plot and the assignment of the plot to an association will be shown. If the plot is not classified, no association will be listed. Occasionally, a&nbsp;plot can be assigned to more than one association. If the full version of the expert system is used, the table will contain a list of associations ranked by decreasing similarity to the plot.</li> <li>If multiple plots are classified, association codes will appear in the header of those plots that were assigned based on the formal definitions. The legends for the codes can be found in a printed version of the&nbsp;<em>Vegetation of the Czech Republic</em>, in its online version (<a href="https://pladias.cz/en/vegetation/">https://pladias.cz/en/vegetation/</a>) or in the expert system txt file. Plots not assigned to any association will be marked with ? and those assigned to more than one association will be marked with +. If the full version of the expert system is used, the output will contain the most similar associations for those plots which have remained unassigned to associations or were assigned to more than one association.</li> <li>Using the formal definition, the expert system normally assigns some plots of a single vegetation stand to a certain association while other plots of the same stand remain unassigned. This means that the stand consists of patches with species composition typical of the given association and patches with a less typical species composition. If different plots from a single relatively homogeneous stand are assigned to different associations, it is appropriate to interpret the stand as transitional between these associations.</li> </ol> <p>This electronic publication of CzechVeg-ESy was supported by the Czech Science Foundation (grant no.&nbsp;17-15168S).</p> <p>***************************************************************************************************************</p> <p><strong>CzechVeg-ESy</strong> je expertn&iacute; syst&eacute;m pro automatickou klasifikaci fytocenologick&yacute;ch sn&iacute;mků z Česk&eacute; republiky do vegetačn&iacute;ch typů definovan&yacute;ch v monografii&nbsp;<em>Vegetace Česk&eacute; republiky&nbsp;</em>(<a href="https://www.sci.muni.cz/botany/vegsci/vegetace.php?lang=en&amp;page=monograph">Chytr&yacute; 2007-2013</a>). Existuj&iacute; dvě hlavn&iacute; verze tohoto expertn&iacute;ho syst&eacute;mu: <strong>Hlavn&iacute; verze 1 (v1) </strong>je origin&aacute;ln&iacute; verze použit&aacute; pro klasifikaci v t&eacute;to n&aacute;rodn&iacute; vegetačn&iacute; monografii, kter&aacute; je podrobně popsan&aacute; v jej&iacute; metodick&eacute; kapitole. Jej&iacute; d&iacute;lč&iacute; verze (označen&eacute; datem) obsahuj&iacute; opravy drobn&yacute;ch chyb a nomenklatury druhů. Tato verze klasifikuje fytocenologick&eacute; sn&iacute;mky pouze do fytocenologick&yacute;ch asociac&iacute;. <strong>Hlavn&iacute; verze 2 (v2)</strong>&nbsp;použ&iacute;v&aacute; stejn&yacute; klasifikačn&iacute; syst&eacute;m, jak&yacute; byl použit ve Vegetaci&nbsp;Česk&eacute; republiky, ale použ&iacute;v&aacute; pokročilej&scaron;&iacute; funkce umožňuj&iacute;c&iacute; přesněj&scaron;&iacute; klasifikaci. Tato verze tak&eacute; prov&aacute;d&iacute; hierarchickou klasifikaci nejen do asociac&iacute;, ale tak&eacute; do svazů a tř&iacute;d (pro sn&iacute;mky nezařazen&eacute; do niž&scaron;&iacute;ch jednotek).</p> <p>Každ&aacute; verze m&aacute; dvě varianty. <strong>Z&aacute;kladn&iacute; varianta&nbsp;</strong>(soubor CzechVeg-ESy-basic-v1.txt) přiřazuje fytocenologick&eacute; sn&iacute;mky do asociac&iacute; na z&aacute;kladě jejich form&aacute;ln&iacute;ch definic vytvořen&yacute;ch metodou Cocktail (<a href="https://doi.org/10.2307/3236796">Bruelheide 2000</a>) v &uacute;pravě podle pr&aacute;ce <a href="https://www.sci.muni.cz/botany/chytry/Koci_etal2003_JVS.pdf">Koč&iacute; et al (2003)</a>. Tyto definice jsou založeny na prezenci sociologick&yacute;ch skupin druhů a dominanci vybran&yacute;ch druhů. Expertn&iacute; syst&eacute;m vyhodnocuje každ&yacute; jednotliv&yacute; fytocenologick&yacute; sn&iacute;mek v datov&eacute;m souboru a řad&iacute; jej do asociace. Klasifikace pomoc&iacute; z&aacute;kladn&iacute; varianty je v&yacute;razně rychlej&scaron;&iacute; než klasifikace pomoc&iacute; pln&eacute; varianty. <strong>Pln&aacute; varianta&nbsp;</strong>(soubor CzechVeg-ESy-full-v1.txt) zaji&scaron;ťuje stejn&eacute; funkce jako z&aacute;kladn&iacute; varianta, ale nav&iacute;c klasifikuje i sn&iacute;mky neklasifikovan&eacute; form&aacute;ln&iacute;mi definicemi, a to na z&aacute;kladě jejich numerick&eacute; podobnosti ke skupin&aacute;m sn&iacute;mků, kter&eacute; vyhovuj&iacute; podm&iacute;nk&aacute;m form&aacute;ln&iacute;ch definic asociac&iacute;. Vět&scaron;inu takov&yacute;ch sn&iacute;mků lze považovat z fytocenologick&eacute;ho hlediska za netypick&eacute; porosty, tj. takov&eacute;, kter&eacute; plně neodpov&iacute;daj&iacute; definovan&yacute;m vegetačn&iacute;m typům, zpravidla kvůli absenci ekologicky specializovan&yacute;ch druhů. Klasifikace na z&aacute;kladě podobnosti se poč&iacute;t&aacute; pomoc&iacute; indexu FPFI (Frequency-Positive Fidelity Index), kter&yacute; definovali&nbsp;<a href="https://www.sci.muni.cz/botany/chytry/Koci_etal2003_JVS.pdf">Koč&iacute; et al. (2003)</a> a&nbsp;<a href="https://doi.org/10.1007/s11258-004-5798-8">Tich&yacute; (2005)</a>.</p> <ol> <li>Fytocenologick&eacute; sn&iacute;mky určen&eacute; k&nbsp;anal&yacute;ze mus&iacute; b&yacute;t uloženy v datab&aacute;zi v&nbsp;programu TURBOVEG 2 (<a href="https://www.synbiosys.alterra.nl/turboveg/">https://www.synbiosys.alterra.nl/turboveg/</a>)&nbsp;s druhov&yacute;m seznamem Czechia-Slovakia-2015 (soubor TurbovegSlBackup_Czechia_slovakia_2015.zip), kter&yacute; je kompatibiln&iacute; s druhov&yacute;mi seznamy Czechia-Slovakia-2012 a Central-Europe.</li> <li>Sn&iacute;mky se exportuj&iacute; z programu&nbsp;TURBOVEG 2 do souboru CC! (Export / Other formats / JUICE input files) a tento soubor se importuje do programu JUICE (<a href="https://www.sci.muni.cz/botany/juice/">https://www.sci.muni.cz/botany/juice/</a>) s použit&iacute;m druhov&eacute;ho seznamu v souboru Checklist-Danihelka-et-al-2012-ver-2019-07-06.txt, č&iacute;mž se nomenklatura konvertuje do nomenklatury odpov&iacute;daj&iacute;c&iacute; Seznamu c&eacute;vnat&yacute;ch rostlin květeny Česk&eacute; republiky (<a href="http://www.preslia.cz/P123Danihelka.pdf">Danihelka et al. 2012</a>).</li> <li>V&nbsp;programu JUICE se zvol&iacute; menu Analysis / Expert system.</li> <li>Tlač&iacute;tkem Load ES File se nahraje do paměti soubor s&nbsp;př&iacute;slu&scaron;n&yacute;m expertn&iacute;m syst&eacute;mem (buď CzechVeg-ESy-basic-v1.txt, nebo&nbsp;CzechVeg-ESy-full-v1.txt).</li> <li>Tlač&iacute;tkem Modify Species Names se uprav&iacute; nomenklatura druhů tak, aby odpov&iacute;dala nomenklatuře použ&iacute;van&eacute; expertn&iacute;m syst&eacute;mem.</li> <li>Obsahuj&iacute;-li sn&iacute;mky určen&eacute; k&nbsp;anal&yacute;ze juveniln&iacute; dřeviny v&nbsp;bylinn&eacute;m patru, je potřeba je vymazat tlač&iacute;tkem Delete Juveniles.</li> <li>Při převodu nomenklatury se v&nbsp;někter&yacute;ch př&iacute;padech převedlo už&scaron;&iacute; pojet&iacute; druhů na &scaron;ir&scaron;&iacute;, č&iacute;mž vznikly v&nbsp;tabulce druhov&eacute; &uacute;daje veden&eacute; pod stejn&yacute;mi jm&eacute;ny. Ty je potřeba sloučit tlač&iacute;tkem Merge Same Spec. Names. Přitom se slouč&iacute; i &uacute;daje stejn&eacute;ho druhu v&nbsp;různ&yacute;ch patrech, protože expertn&iacute; syst&eacute;m předpokl&aacute;d&aacute; jen jeden v&yacute;skyt stejn&eacute;ho druhov&eacute;ho jm&eacute;na v&nbsp;jednom sn&iacute;mku.</li> <li>Pokud je použ&iacute;v&aacute;na pln&aacute; (Full) verze expertn&iacute;ho syst&eacute;mu, je potřeba v&nbsp;ok&eacute;nku vpravo dole nastavit prahovou hodnotu podobnosti pro přiřazov&aacute;n&iacute; sn&iacute;mků k&nbsp;asociac&iacute;m pomoc&iacute; podobnosti. Č&iacute;m vy&scaron;&scaron;&iacute; hodnota, t&iacute;m m&eacute;ně sn&iacute;mků se přiřad&iacute;, ale přiřad&iacute; se ty, kter&eacute; se dan&eacute; asociaci v&iacute;ce podobaj&iacute;. Při hodnotě 0 se přiřad&iacute; v&scaron;echny sn&iacute;mky, ale někter&eacute; budou dan&eacute; asociaci velmi nepodobn&eacute;.</li> <li>Spust&iacute; se běh expertn&iacute;ho syst&eacute;mu, a to buď tlač&iacute;tkem Classify Relev&eacute; (bude se klasifikovat jeden sn&iacute;mek, na kter&yacute; se předt&iacute;m kliklo my&scaron;&iacute;) nebo Classify [colour] Relev&eacute;s (budou se klasifikovat v&scaron;echny sn&iacute;mky vybran&eacute; barvy).</li> <li>Při klasifikaci jednoho sn&iacute;mku se zobraz&iacute; v&nbsp;tabulce druhov&eacute; skupiny, jejich zastoupen&iacute; v&nbsp;dan&eacute;m sn&iacute;mku a asociace, do kter&eacute; byl sn&iacute;mek přiřazen pomoc&iacute; form&aacute;ln&iacute; definice. Pokud přiřazen nebyl, nezobraz&iacute; se ž&aacute;dn&aacute; asociace.&nbsp;Sn&iacute;mek může b&yacute;t přiřazen i do v&iacute;ce než jedn&eacute; asociace. Při použit&iacute; pln&eacute; verze expertn&iacute;ho syst&eacute;mu se do tabulky vyp&iacute;&scaron;&iacute; asociace v&nbsp;pořad&iacute; klesaj&iacute;c&iacute; podobnosti ke sn&iacute;mku, a to u těch sn&iacute;mků, kter&eacute; nebyly přiřazeny do ž&aacute;dn&eacute; asociace nebo byly přiřazeny do v&iacute;ce než jedn&eacute; asociace.</li> <li>Při klasifikaci v&iacute;ce sn&iacute;mků se do z&aacute;hlav&iacute; tabulky vep&iacute;&scaron;&iacute; k&oacute;dy asociac&iacute; u těch sn&iacute;mků, kter&eacute; se přiřadily na z&aacute;kladě form&aacute;ln&iacute;ch definic. Převod k&oacute;dů na jm&eacute;na asociac&iacute; lze dohledat v&nbsp;ti&scaron;těn&eacute; verzi&nbsp;<em>Vegetace Česk&eacute; republiky</em>, v jej&iacute; online verzi (<a href="https://pladias.cz/en/vegetation/">https://pladias.cz/vegetation/</a>) nebo v&nbsp;textov&eacute;m souboru expertn&iacute;ho syst&eacute;mu. U sn&iacute;mků, kter&eacute; se nepřiřadily k&nbsp;ž&aacute;dn&eacute; asociaci, se v z&aacute;hlav&iacute; zobraz&iacute; znak ?. U sn&iacute;mků přiřazen&yacute;ch do v&iacute;ce než jedn&eacute; asociace se zobraz&iacute; znak +. Při použit&iacute; pln&eacute; verze expertn&iacute;ho syst&eacute;mu se do tabulky vyp&iacute;&scaron;&iacute; nejpodobněj&scaron;&iacute; asociace u těch sn&iacute;mků, kter&eacute; nebyly přiřazeny do ž&aacute;dn&eacute; asociace nebo byly přiřazeny do v&iacute;ce než jedn&eacute; asociace.</li> <li>Expertn&iacute; syst&eacute;m běžně přiřazuje pomoc&iacute; form&aacute;ln&iacute;ch definic někter&eacute; sn&iacute;mky v&nbsp;porostu nebo lok&aacute;lně rozli&scaron;ovan&eacute;m rostlinn&eacute;m společenstvu do určit&eacute; asociace a jin&eacute; do ž&aacute;dn&eacute; asociace, což znamen&aacute;, že se porost skl&aacute;d&aacute; z&nbsp;m&iacute;st s&nbsp;druhov&yacute;m složen&iacute;m typick&yacute;m pro danou asociaci a m&iacute;st s&nbsp;m&eacute;ně typick&yacute;m druhov&yacute;m složen&iacute;m. Pokud expertn&iacute; syst&eacute;m přiřad&iacute; různ&eacute; sn&iacute;mky z&nbsp;jednoho relativně homogenn&iacute;ho porostu k&nbsp;různ&yacute;m asociac&iacute;m, je vhodn&eacute; porost interpretovat jako přechodn&yacute; mezi těmito asociacemi.</li> </ol> <p>Tato elektronick&aacute; publikace expertn&iacute;ho syst&eacute;mu CzechVeg-ESy byla podpořena Grantovou agenturou Česk&eacute; republiky (grant 17-15168S).</p>

opencc-by-4.0Jan 2020View details →
zenodo48/100

U2- and U12-type intron classifications for Physarum polycephalum in BED format

<p>A BED file containing intron information for <em>Physarum polycephalum</em>&nbsp;introns classified as U2- or U12-type by <a href="https://github.com/glarue/intronIC">intronIC</a>. This data is associated with the following manuscript:&nbsp;https://doi.org/10.1101/2020.10.12.336362; the genome and annotation file&nbsp;used to identify the introns are available here:&nbsp;https://doi.org/10.5281/zenodo.4086119.</p> <p>&nbsp;</p> <p>The file columns are:</p> <p>1. Genome FASTA record name (scaffold)</p> <p>2. Intron start coordinate (0-indexed)</p> <p>3. Intron end coordinate (1-indexed)</p> <p>4. Intron label from intronIC</p> <p>5. U12-type probability score (0-100); introns with scores &gt; 95 were considered U12-type in the manuscript</p> <p>6. Strand</p>

opencc-by-4.0Oct 2020View details →
zenodo48/100

Ground Truth and Automated Classification from Copernicus Sentinel-2 Imagery

<p>Ground-Truth and Sentinel2 imagery classification of <em>Trees Outside Forest</em> in an agroforestry landscape in Umbria,&nbsp;Italy.</p> <p>Location:&nbsp;Alfina plains, Castelgiorgio area, Umbria, Italy.&nbsp;Reference system:&nbsp;EPSG:32632&nbsp;(WGS84, UTM zone 32 North)&nbsp;Extent: West 740609 &mdash; East 750828,&nbsp;South 4726490 &mdash; North 4737250</p> <p>Dataset&nbsp;format: geopackage, a single file&nbsp;<strong>data.gpkg</strong>&nbsp;containing 9 vector layers (in alphabetical order):</p> <ol> <li>Areas&nbsp;&mdash; Areas of interest, 2 polygons</li> <li>Classification&nbsp;&mdash; Automated classification from Sentinel2 imagery, 11781 polygons</li> <li>Hedgerows1&nbsp;&mdash; Ground truth, hedgerows of Area1, 148 lines</li> <li>Hedgerows2&nbsp;&mdash; Ground truth, hedgerows of Area2, 135 lines</li> <li>Sentinel2&nbsp;&mdash; Sentinel2 scenes footprint, one&nbsp;polygon</li> <li>Trees1&nbsp;&mdash; Ground truth, isolated trees of Area1, 55 points</li> <li>Trees2&nbsp;&mdash; Ground truth, isolated trees of Area2, 64 points</li> <li>Woods1&nbsp;&mdash; Ground truth, small forest patches of Area1, 33 polygons</li> <li>Woods2&nbsp;&mdash; Ground truth, small forest patches of Area2, 37 polygons</li> </ol> <p>Accompanying map:&nbsp;<strong>map.qgz</strong>, Qgis 3.6 format. The geopackage&nbsp;dataset is supposed to be stored in the same directory of the map (relative path = ./)</p> <p>Dataset description and metadata: <strong>meta.pdf</strong>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2019View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record