Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
2,609
datasets available to search
ShareScore release 0.9.0
Dataset results
2,609 results for “Web”
Food Web of Sarracenia Purpurea in United States and Canada 1999-2011
How food webs are structured and how their structure and dynamics vary through time and space is a central focus of research in community ecology. We documented structural variation in the aquatic food web inhabiting pitcher-shaped leaves of the carnivorous pitcher plant Sarracenia purpurea across the geographic range of the plant (from Florida north to Labrador and west to British Columbia); examined temporal variation in this food web with detailed experiments in Massachusetts and Vermont; experimentally manipulated top-down and bottom-up processes in this food web in Massachusetts; and developed a dynamic simulation model of this food web that incorporates metacommunity dynamics.
Datos brutos de mensajes publicados en usuarios vinculados a medios informativos digitales españoles en X, FB y Web
<p>Datos brutos recabados en el marco del proyecto Hatemedia (PID2020-114584GB-I00), financiado por <strong>MCIN/ AEI /10.13039/501100011033.</strong></p> <p>Los datos recopilados corresponden a cinco medios informativos digitales, tomados como casos de estudio en el proyecto al que corresponde estos datos, tanto en la plataforma X, como en FB y en sus portales web; durante 2021-2022.</p> <p>El proceso de recolección de estos datos fueron recolectados, a partir del procedimiento expuesto en:</p> <p>Ruiz-Iniesta, Almudena; Blanco Valencia, Xiomara; Pérez, Daniel; De Gregorio Vicente, Oscar; José Cubillas, Juan; Julio Montero Díaz, https://orcid.org/0000-0002-4145-7424; et al. (2024). Informe final sobre el scrapeo de datos brutos obtenidos en medios informativos digitales españoles en X, FB y Web. figshare. Presentation. https://doi.org/10.6084/m9.figshare.25187591.v2</p> <p>Los datos fueron recolectados con el apoyo de la empresa colaboradora: <a href="https://www.possibleinc.com/?lang=en">Possible Inc</a> </p> <p>Para más información del proyecto, ingresar en: https://www.hatemedia.es/</p> <p>Si deseas tener acceso a estos datos, contactar con los IPs del proyecto, a traves de: elias.said@unir.net, en copia a julio.montero@unir.net</p>
Modeling Foundation Species in Food Webs
Foundation species are basal species that play an important role in determining community composition by physically structuring ecosystems and modulating ecosystem processes. Foundation species largely operate via non-trophic interactions, presenting a challenge to incorporating them into food-web models. Here, we used non-linear, bioenergetic predator-prey models to explore the role of foundation species and their non-trophic effects. We explored four types of models in which the foundation species reduced the metabolic rates of species in a specific trophic position. We examined the outcomes of each of these models for six metabolic rate “treatments” in which the foundation species altered the metabolic rates of associated species by one-tenth to ten times their allometric baseline metabolic rates. For each model simulation, we looked at how foundation species influenced food-web structure during community assembly and the subsequent change in food-web structure when the foundation species was removed. When a foundation species lowered the metabolic rate of only basal species the resultant webs were complex, species-rich, and robust to foundation species removals. On the other hand, when a foundation species lowered the metabolic rate of only consumer species, all species, or no species the resultant webs were species poor and the subsequent removal of the foundation species webs resulted in the further loss of species and complexity. This suggests that in nature we should look for foundation species to predominantly facilitate basal species.
Food Web Isotope Study at North Temperate Lakes LTER 2004 - 2010
Stable isotope ratios (d13C and d15N) were measured in archived scale samples of seven fish species captured in Sparkling Lake over the period 1981 through 2009. Stable isotopes of benthic and pelagic food web end members (macroinvertebrates and zooplankton, respectively) were also measured. Zooplankton were from samples taken in Sparkling Lake, Allequash Lake, and Big Muskellunge Lake. Zoobenthos were from samples taken in Sparkling Lake. Isotope Study abstract paragraph 2 Number of sites: 3
Cascade Project at North Temperate Lakes LTER High Frequency Sonde Data from Food Web Resilience Experiment 2008 - 2011
High-frequency sonde data collected from the surface waters of two lakes in Upper Peninsula of Michigan during the summers of 2008-2011. The food web of Peter Lake was slowly transformed by gradual additions of Largemouth bass (Micropterus salmoides) while Paul Lake was an unmanipulated reference. Sonde data were used to calculate resilience indicators to evaluate the stability of the food web and to calculate ecosystem metabolism.
Cebulka (Polish dark web cryptomarket and image board) messages data
<h3><strong>General Information</strong></h3> <p>1. <strong>Title of Dataset</strong></p> <p>Cebulka (Polish dark web cryptomarket and image board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Patrycja Cheba (Jagiellonian University); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: “Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade” (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4.<strong> Data Source</strong></p> <p>Polish dark web cryptomarket and image board called Cebulka (<a href="http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php">http://cebulka7uxchnbpvmqapg5pfos4ngaxglsktzvha7a5rigndghvadeyd.onion/index.php</a>). </p> <p>5. <strong>Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project focuses on studying internet behavior concerning disruptive actions, particularly emphasizing the online narcotics market in Poland. The research seeks to (1) investigate how the open internet, including social media, is used in the drug trade; (2) outline the significance of darknet platforms in the distribution of drugs; and (3) explore the complex exchange of content related to the drug trade between the surface web and the darknet, along with understanding meanings constructed within the drug subculture.</p> <p>Within this context, Cebulka is identified as a critical digital venue in Poland’s dark web illicit substances scene. Besides serving as a marketplace, it plays a crucial role in shaping the narratives and discussions prevalent in the drug subculture. The dataset has proved to be a valuable tool for performing the analyses needed to achieve the project’s objectives.</p> <h3><strong>Data Content</strong></h3> <p>6. <strong>Data Description</strong></p> <p>The data was collected in three periods, i.e., in January 2023, June 2023, and January 2024.</p> <p>The dataset comprises a sample of messages posted on Cebulka from its inception until January 2024 (including all the messages with drug advertisements). These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. The dataset is organized into two directories. The “cebulka_adverts” directory contains posts related to drug advertisements (both advertisements and comments). In contrast, the “cebulka_community” directory holds a sample of posts from other parts of the cryptomarket, i.e., those not related directly to trading drugs but rather focusing on discussing illicit substances. The dataset consists of 16,842 posts.</p> <p>7. <strong>Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>8. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following files:</p> <ul> <li>Zipped .txt files (“cebulka_adverts.zip” and “cebulka_community.zip”) containing all messages. These files are organized into individual directories that mirror the folder structure found on Cebulka.</li> <li>Two .csv files that list all the messages, including file names and the content of each post. The first .csv lists messages from “cebulka_adverts.zip,” and the second .csv lists messages from “cebulka_community.zip.”</li> </ul> <h3><strong>Ethical Considerations</strong></h3> <p>9. <strong>Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the messages, utilizing automated systems for irreversible hashing. Recognizing that automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>
Hyperreal Talk (Polish clear web message board) messages data
<h3><strong>General Information</strong></h3> <p>1.<strong> Title of Dataset</strong></p> <p>Hyperreal Talk (Polish clear web message board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: “Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade” (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4.<strong> Data Source</strong></p> <p>Polish clear web message board called Hyperreal Talk (<a href="https://hyperreal.info/talk/">https://hyperreal.info/talk/</a>).</p> <p>5.<strong> Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project delves into internet dynamics within disruptive activities, specifically focusing on the online drug trade in Poland. It aims to (1) examine the utilization of the open internet, including social media, in the drug trade; (2) delineate the role of darknet environments in narcotics distribution; and (3) uncover the intricate flow of drug trade-related content and its meanings between the open web and the darknet, and how these meanings are shaped within the so-called drug subculture.</p> <p>The Hyperreal Talk forum emerges as a pivotal online space on the Polish internet, serving as a hub for discussions and the exchange of knowledge and experiences concerning drug use. It plays a crucial role in investigating the narratives and discourses that shape the drug subculture and the broader societal perceptions of drug consumption. The dataset has been instrumental in conducting analyses pertinent to the earlier project goals.</p> <p>6. <strong>Collection Method</strong></p> <p>The dataset was compiled using the Scrapy framework, a web crawling and scraping library for Python. This tool facilitated systematic content extraction from the targeted message board.</p> <p>7.<strong> Collection Date</strong></p> <p>The data was collected in two periods, i.e., in September 2023 and November 2023.</p> <h3><strong>Data Content</strong></h3> <p>8. <strong>Data Description</strong></p> <p>The dataset comprises all messages posted on the Polish-language Hyperreal Talk message board from its inception until November 2023. These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. The dataset is organized into two directories: “hyperreal” and “hyperreal_hidden.” The “hyperreal” directory contains accessible posts without needing to log in to Hyperreal Talk, while the “hyperreal_hidden” directory holds posts that can only be viewed by logged-in users. For each directory, a .txt file has been prepared detailing the structure of the message board folders from which the posts were extracted. The dataset includes 6,248,842 posts.</p> <p>9.<strong> Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>10. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following files:</p> <ul> <li>Zipped .txt files (hyperreal.zip) containing messages that are visible without logging into Hyperreal Talk. These files are organized into individual directories that mirror the folder structure found on the Hyperreal Talk message board.</li> <li>Zipped .txt files (hyperreal_hidden.zip) containing messages that are visible only after logging into Hyperreal Talk. Similar to the first type, these files are organized into directories corresponding to the website’s folder structure.</li> <li>A .csv file that lists all the messages, including file names and the content of each post.</li> </ul> <h3><strong>Accessibility and Usage</strong></h3> <p>11.<strong> Access Conditions</strong></p> <p>The data can be accessed without any restrictions.</p> <p>12. <strong>Related Documentation</strong></p> <p>Attached are .txt files detailing the tree of folders for “hyperreal.zip” and “hyperreal_hidden.zip.”</p> <p>Documentation on the Python regular expressions used for scraping, cleaning, processing, and anonymizing the data can be found on GitHub at the following URLs:</p> <ul> <li><a href="https://github.com/LeszekSwieca/Project_2021-43-B-HS6-00710">https://github.com/LeszekSwieca/Project_2021-43-B-HS6-00710</a></li> <li><a href="https://github.com/HaitaoShi/Scrapy_hyperreal">https://github.com/HaitaoShi/Scrapy_hyperreal</a>"</li> </ul> <h3><strong>Ethical Considerations</strong></h3> <p>13. <strong>Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the messages, utilizing automated systems for irreversible hashing. Recognizing that scraping and automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>
Dopek.eu (Polish clear web and dark web message board) messages data
<h3><strong>General Information</strong></h3> <p>1. <strong>Title of Dataset</strong></p> <p>Dopek.eu (Polish clear web and dark web message board) messages data.</p> <p>2. <strong>Data Collectors</strong></p> <p>Haitao Shi (The University of Edinburgh, UK); Leszek Świeca (Kazimierz Wielki University in Bydgoszcz, Poland).</p> <p>3. <strong>Funding Information</strong></p> <p>The dataset is part of the research supported by the Polish National Science Centre (Narodowe Centrum Nauki) grant 2021/43/B/HS6/00710.</p> <p>Project title: “Rhizomatic networks, circulation of meanings and contents, and offline contexts of online drug trade” (2022-2025; PLN 956 620; funding institution: Polish National Science Centre [NCN], call: OPUS 22; Principal Investigator: Piotr Siuda [Kazimierz Wielki University in Bydgoszcz, Poland]).</p> <h3><strong>Data Collection Context</strong></h3> <p>4. <strong>Data Source</strong></p> <p>Clear web and dark web message board called dopek.eu (<a href="https://dopek.eu/">https://dopek.eu/</a>). </p> <p>5.<strong> Purpose</strong></p> <p>This dataset was developed within the abovementioned project. The project delves into internet dynamics within disruptive activities, specifically focusing on the online drug trade in Poland. It aims to (1) examine the utilization of the open internet, including social media, in the drug trade; (2) delineate the role of darknet environments in narcotics distribution; and (3) uncover the intricate flow of drug trade-related content and its meanings between the open web and the darknet, and how these meanings are shaped within the so-called drug subculture.</p> <p>The dopek.eu forum emerges as a pivotal online space on the Polish internet, serving as a hub for trading, discussions, and the exchange of knowledge and experiences concerning the use of the so-called new psychoactive substances (designer drugs). The dataset has been instrumental in conducting analyses pertinent to the earlier project goals.</p> <p>6.<strong> Collection Method</strong></p> <p>The dataset was compiled using the Scrapy framework, a web crawling and scraping library for Python. This tool facilitated systematic content extraction from the targeted message board.</p> <p>7.<strong> Collection Date</strong></p> <p>The data was collected in October 2023.</p> <h3><strong>Data Content</strong></h3> <p>8.<strong> Data Description</strong></p> <p>The dataset comprises all messages posted on dopek.eu from its inception until October 2023. These messages include the initial posts that start each thread and the subsequent posts (replies) within those threads. A .txt file has been prepared detailing the structure of the message board folders from which the posts were extracted. The dataset includes 171,121 posts.</p> <p>9.<strong> Data Cleaning, Processing, and Anonymization</strong></p> <p>The data has been cleaned and processed using regular expressions in Python. Additionally, all personal information was removed through regular expressions. The data has been hashed to exclude all identifiers related to instant messaging apps and email addresses. Furthermore, all usernames appearing in messages have been eliminated.</p> <p>10. <strong>File Formats and Variables/Fields</strong></p> <p>The dataset consists of the following types of files:</p> <ul> <li>Zipped .txt files (dopek.zip) containing all messages (posts).</li> <li>A .csv file that lists all the messages, including file names and the content of each post.</li> </ul> <h3><strong>Accessibility and Usage</strong></h3> <p><strong>11. Access Conditions</strong></p> <p>The data can be accessed without any restrictions.</p> <p><strong>12. Related Documentation</strong></p> <p>Attached are .txt files detailing the tree of folders for “dopek.zip”.</p> <h3><strong>Ethical Considerations</strong></h3> <p><strong>13. Ethics Statement</strong></p> <p>A set of data handling policies aimed at ensuring safety and ethics has been outlined in the following paper:</p> <p>Harviainen, J.T., Haasio, A., Ruokolainen, T., Hassan, L., Siuda, P., Hamari, J. (2021). Information Protection in Dark Web Drug Markets Research [in:] Proceedings of the 54th Hawaii International Conference on System Sciences, HICSS 2021, Grand Hyatt Kauai, Hawaii, USA, 4-8 January 2021, Maui, Hawaii, (ed.) Tung X. Bui, Honolulu, HI, pp. 4673-4680.</p> <p>The primary safeguard was the early-stage hashing of usernames and identifiers from the posts, utilizing automated systems for irreversible hashing. Recognizing that scraping and automatic name removal might not catch all identifiers, the data underwent manual review to ensure compliance with research ethics and thorough anonymization.</p>
Classification of web-based Digital Humanities projects leveraging information visualisation techniques
<h2>Description</h2> <p>This dataset contains a list of 186 Digital Humanities projects leveraging information visualisation methods. Each project has been classified according to visualisation and interaction techniques, narrativity and narrative solutions, domain, methods for the representation of uncertainty and interpretation, and the employment of critical and custom approaches to visually represent humanities data.</p> <p> </p> <h2>Classification schema: categories and columns</h2> <p>The <code>project_id</code> column contains unique internal identifiers assigned to each project. Meanwhile, the <code>last_access</code> column records the most recent date (in DD/MM/YYYY format) on which each project was reviewed based on the web address specified in the <code>url</code> column.<br>The remaining columns can be grouped into descriptive categories aimed at characterising projects according to different aspects:</p> <p> </p> <p><strong>Narrativity.</strong> It reports the presence of information visualisation techniques employed within narrative structures. Here, the term narrative encompasses both author-driven linear data stories and more user-directed experiences where the narrative sequence is determined by user exploration [1]. We define 2 columns to identify projects using visualisation techniques in narrative, or non-narrative sections. Both conditions can be true for projects employing visualisations in both contexts. Columns:</p> <ul> <li> <p><code>non_narrative</code> (boolean)</p> </li> <li> <p><code>narrative</code> (boolean)</p> </li> </ul> <p> </p> <p><strong>Domain.</strong> The humanities domain to which the project is related. We rely on [2] and the chapters of the first part of [3] to abstract a set of general domains. Column:</p> <ul> <li> <p><code>domain</code> (categorical):</p> </li> <ul> <li> <p>History and archaeology</p> </li> <li> <p>Art and art history</p> </li> <li> <p>Language and literature</p> </li> <li> <p>Music and musicology</p> </li> <li> <p>Multimedia and performing arts</p> </li> <li> <p>Philosophy and religion</p> </li> <li> <p>Other: both extra-list domains and cases of collections without a unique or specific thematic focus.</p> </li> </ul> </ul> <p> </p> <p><strong>Visualisation of uncertainty and interpretation.</strong> Buiding upon the frameworks proposed by [4] and [5], a set of categories was identified, highlighting a distinction between precise and impressional communication of uncertainty. Precise methods explicitly represent quantifiable uncertainty such as missing, unknown, or uncertain data, precisely locating and categorising it using visual variables and positioning. Two sub-categories are interactive distinction, when uncertain data is not visually distinguishable from the rest of the data but can be dynamically isolated or included/excluded categorically through interaction techniques (usually filters); and visual distinction, when uncertainty visually “emerges” from the representation by means of dedicated glyphs and spatial or visual cues and variables. On the other hand, impressional methods communicate the constructed and situated nature of data [6], exposing the interpretative layer of the visualisation and indicating more abstract and unquantifiable uncertainty using graphical aids or interpretative metrics. Two sub-categories are: ambiguation, when the use of graphical expedients—like permeable glyph boundaries or broken lines—visually convey the ambiguity of a phenomenon; and interpretative metrics, when expressive, non-scientific, or non-punctual metrics are used to build a visualisation. Column:</p> <ul> <li> <p><code>uncertainty_interpretation</code> (categorical):</p> </li> <ul> <li> <p>Interactive distinction</p> </li> <li> <p>Visual distinction</p> </li> <li> <p>Ambiguation</p> </li> <li> <p>Interpretative metrics</p> </li> </ul> </ul> <p> </p> <p><strong>Critical adaptation.</strong> We identify projects in which, with regards to at least a visualisation, the following criteria are fulfilled: 1) avoid repurposing of prepackaged, generic-use, or ready-made solutions; 2) being tailored and unique to reflect the peculiarities of the phenomena at hand; 3) avoid simplifications to embrace and depict complexity, promoting time-consuming visualisation-based inquiry. Column:</p> <ul> <li> <p><code>critical_adaptation</code> (boolean)</p> </li> </ul> <p> </p> <p><strong>Non-temporal visualisation techniques.</strong> We adopt and partially adapt the terminology and definitions from [7]. A column is defined for each type of visualisation and accounts for its presence within a project, also including stacked layouts and more complex variations. Columns and inclusion criteria:</p> <ul> <li> <p><code>plot</code> (boolean): visual representations that map data points onto a two-dimensional coordinate system.</p> </li> <li> <p><code>cluster_or_set</code> (boolean): sets or cluster-based visualisations used to unveil possible inter-object similarities.</p> </li> <li> <p><code>map</code> (boolean): geographical maps used to show spatial insights. While we do not specify the variants of maps (e.g., pin maps, dot density maps, flow maps, etc.), we make an exception for maps where each data point is represented by another visualisation (e.g., a map where each data point is a pie chart) by accounting for the presence of both in their respective columns.</p> </li> <li> <p><code>network</code> (boolean): visual representations highlighting relational aspects through nodes connected by links or edges.</p> </li> <li> <p><code>hierarchical_diagram</code> (boolean): tree-like structures such as tree diagrams, radial trees, but also dendrograms. They differ from networks for their strictly hierarchical structure and absence of closed connection loops.</p> </li> <li> <p><code>treemap</code> (boolean): still hierarchical, but highlighting quantities expressed by means of area size. It also includes circle packing variants.</p> </li> <li> <p><code>word_cloud</code> (boolean): clouds of words, where each instance’s size is proportional to its frequency in a related context</p> </li> <li> <p><code>bars</code> (boolean): includes bar charts, histograms, and variants. It coincides with “bar charts” in [7] but with a more generic term to refer to all bar-based visualisations.</p> </li> <li> <p><code>line_chart</code> (boolean): the display of information as sequential data points connected by straight-line segments.</p> </li> <li> <p><code>area_chart</code> (boolean): similar to a line chart but with a filled area below the segments. It also includes density plots.</p> </li> <li> <p><code>pie_chart</code> (boolean): circular graphs divided into slices which can also use multi-level solutions.</p> </li> <li> <p><code>plot_3d</code> (boolean): plots that use a third dimension to encode an additional variable.</p> </li> <li> <p><code>proportional_area</code> (boolean): representations used to compare values through area size. Typically, using circle- or square-like shapes.</p> </li> <li> <p><code>other</code> (boolean): it includes all other types of non-temporal visualisations that do not fall into the aforementioned categories.</p> </li> </ul> <p> </p> <p><strong>Temporal visualisations and encodings.</strong> In addition to non-temporal visualisations, a group of techniques to encode temporality is considered in order to enable comparisons with [7]. Columns:</p> <ul> <li> <p><code>timeline</code> (boolean): the display of a list of data points or spans in chronological order. They include timelines working either with a scale or simply displaying events in sequence. As in [7], we also include structured solutions resembling Gantt chart layouts.</p> </li> </ul> <ul> <li> <p><code>temporal_dimension</code> (boolean): to report when time is mapped to any dimension of a visualisation, with the exclusion of timelines. We use the term “dimension” and not “axis” as in [7] as more appropriate for radial layouts or more complex representational choices.</p> </li> <li> <p><code>animation</code> (boolean): temporality is perceived through an animation changing the visualisation according to time flow.</p> </li> <li> <p><code>visual_variable</code> (boolean): another visual encoding strategy is used to represent any temporality-related variable (e.g., colour).</p> </li> </ul> <p> </p> <p><strong>Interaction techniques.</strong> A set of categories to assess affordable interaction techniques based on the concept of user intent [8] and user-allowed data actions [9]. The following categories roughly match the “processing”, “mapping”, and “presentation” actions from [9] and the manipulative subset of methods of the “how” an interaction is performed in the conception of [10]. Only interactions that affect the visual representation or the aspect of data points, symbols, and glyphs are taken into consideration. Columns:</p> <ul> <li> <p><code>basic_selection</code> (boolean): the demarcation of an element either for the duration of the interaction or more permanently until the occurrence of another selection.</p> </li> <li> <p><code>advanced_selection</code> (boolean): the demarcation involves both the selected element and connected elements within the visualisation or leads to brush and link effects across views. Basic selection is tacitly implied.</p> </li> <li> <p><code>navigation</code> (boolean): interactions that allow moving, zooming, panning, rotating, and scrolling the view but only when applied to the visualisation and not to the web page. It also includes “drill” interactions (to navigate through different levels or portions of data detail, often generating a new view that replaces or accompanies the original) and “expand” interactions generating new perspectives on data by expanding and collapsing nodes.</p> </li> <li> <p><code>arrangement</code> (boolean): methods to organise visualisation elements (symbols, glyphs, etc.) or multi-visualisation layouts spatially through drag and drop or according to a criterion via more automatic triggers.</p> </li> <li> <p><code>change</code> (boolean): visual encoding alterations involving different aspects of visualisation as a whole: the same content is presented with another visualisation technique; the change involves symbols or glyphs aspect (colour, size, shape, etc.); the visualisation type is unaltered, but the layout variant changes (e.g., to stacked layouts); or other changes like axes inversion and scale modifications. The presence of all the visualisation techniques involved in a change is reported.</p> </li> <li> <p><code>visualisation_filter</code> (boolean): filters to exclude or include visualisation elements with respect to defined criteria, without reloading or generating a new visualisation. Unlike options triggering the fetch of new data to alter the visualisation content, filters seamlessly operate on existing visual elements.</p> </li> <li> <p><code>collection_filter</code> (boolean): the interaction with visualised elements acts as a filter for a related collection or list of items (e.g., clicking a region on a map filters a list of items according to spatial metadata).</p> </li> <li> <p><code>aggregation</code> (boolean): changes to the granularity of visual elements according to a variable. It produces either visual data summarisations or segregations.</p> </li> <li> <p><code>btfw_interaction</code> (boolean): to identify the use of “breaking the fourth wall interactions” as defined [11]. It applies only to narratives.</p> </li> </ul> <p> </p> <p><strong>Narrative flow factors.</strong> Other categories aim to identify patterns in the design of narrative solutions. It is worth noticing that a project with multiple and diverse narratives can potentially report multiple design choices for the same column. Part of the factors and definitions from [12] are here re-used and adapted.</p> <p><em>Story layout </em>columns define the layout, or genre, of the narrative format:</p> <ul> <li> <p><code>document_layout</code> (boolean)</p> </li> <li> <p><code>slideshow_layout</code> (boolean)</p> </li> <li> <p><code>hybrid_layout</code> (boolean): mixing document and slideshow layouts.</p> </li> <li> <p><code>other_layout</code> (boolean): more complex solutions.</p> </li> </ul> <p><em>Role of visualisation</em> columns describe the role visualisations detain with respect to the entire story, in particular, with reference to the textual part of the narratives:</p> <ul> <li><code>equal_role</code> (boolean): visualisations and text play an equal role in the narrative.</li> <li><code>figure_role</code> (boolean): visualisations are supporting elements compared to the role of text.</li> <li><code>annotated_role</code> (boolean): visualisations are the drivers of the narrative.</li> </ul> <p><em>Story progression</em> columns categorise the shape of possible story paths:</p> <ul> <li> <p><code>linear_progression</code> (categorical): strongly author-driven or user-directed narrative. Possible values specify the potential to skip certain parts while not having a fully explorative experience:</p> </li> <ul> <li> <p>Skip</p> </li> <li> <p>No-skip</p> </li> </ul> <li> <p><code>user_directed</code> (bool): users can select a path among multiple alternatives and compose narrative pieces, providing a broder degree of interaction and exploration possibilities [1]. If a linear path can be suggested, here it remains merely one option among many others. Differently from a linear-skip approach, it has a low level of guidance oriented towards linear navigation.</p> </li> </ul> <p><em>Navigation input </em>columns define the ways users can move through the narrative:</p> <ul> <li> <p><code>button_input</code> (boolean)</p> </li> <li> <p><code>scroll_input</code> (boolean)</p> </li> <li> <p><code>slider_input</code> (boolean)</p> </li> </ul> <p><em>Navigation progress </em>columns describe methods through which the reader perceives its placement within the narrative:</p> <ul> <li> <p><code>text_progression</code> (boolean): text or numbers act as signifiers for user position.</p> </li> <li> <p><code>dots_progression</code> (boolean)</p> </li> <li> <p><code>visualisation_progression</code> (boolean): the visualisation used in the narrative, or a visualised progress widget acts as a signifier for user position.</p> </li> </ul> <p><em>Level of control </em>columns describe how much control a reader has over the text, visualisations, and animated transitions. Control could be discrete (D) when it triggers the motion, continuous (C) when it can act throughout all the keyframes, or hybrid (H) if it supports aspects of both. When animation is absent, control can be not available (NA). In particular, while visualisation control is related to the visualisation as a whole (e.g., the entire scatter plot moving up or down the page), the animated transition is related to more specific, data-relevant motion.<br>Columns:</p> <ul> <li> <p><code>text_control</code> (categorical):</p> </li> <ul> <li> <p>D</p> </li> <li> <p>C</p> </li> <li> <p>H</p> </li> </ul> <li> <p><code>visualisation_control</code> (categorical):</p> </li> <ul> <li> <p>D</p> </li> <li> <p>C</p> </li> <li> <p>H</p> </li> </ul> <li> <p><code>animation_control</code> (categorical):</p> </li> <ul> <li> <p>D</p> </li> <li> <p>C</p> </li> <li> <p>H</p> </li> <li> <p>NA</p> </li> </ul> </ul> <p> </p> <h2>References</h2> <p>[1] E. Segel and J. Heer, “Narrative Visualization: Telling Stories with Data,” IEEE Trans. Visual. Comput. Graphics, vol. 16, no. 6, pp. 1139–1148, 2010, doi: 10.1109/TVCG.2010.179.</p> <p>[2] M. Terras, J. Nyhan, and E. Vanhoutte, Defining Digital Humanities: A Reader. Routledge, 2016.</p> <p>[3] S. Schreibman, R. G. Siemens, and J. Unsworth, Eds., A companion to digital humanities. in Blackwell companions to literature and culture, no. 26. Malden, MA: Blackwell Pub, 2004.</p> <p>[4] C. Kinkeldey, A. M. MacEachren, and J. Schiewe, “How to Assess Visual Communication of Uncertainty? A Systematic Review of Geospatial Uncertainty Visualisation User Studies,” The Cartographic Journal, vol. 51, no. 4, pp. 372–386, 2014, doi: 10.1179/1743277414Y.0000000099.</p> <p>[5] G. Panagiotidou, H. Lamqaddam, J. Poblome, K. Brosens, K. Verbert, and A. Vande Moere, “Communicating Uncertainty in Digital Humanities Visualization Research,” IEEE Transactions on Visualization and Computer Graphics, vol. 29, no. 1, pp. 635–645, Jan. 2023, doi: 10.1109/TVCG.2022.3209436.</p> <p>[6] J. Drucker, “Humanities Approaches to Graphical Display,” Digital Humanities Quarterly, vol. 5, no. 1, 2011, Accessed: Sep. 17, 2024. [Online]. Available: <a href="https://www.digitalhumanities.org/dhq/vol/5/1/000091/000091.html">https://www.digitalhumanities.org/dhq/vol/5/1/000091/000091.html</a></p> <p>[7] F. Windhager et al., “Visualization of Cultural Heritage Collection Data: State of the Art and Future Challenges,” IEEE Trans. Visual. Comput. Graphics, vol. 25, no. 6, pp. 2311–2330, Jun. 2019, doi: 10.1109/TVCG.2018.2830759.</p> <p>[8] J. S. Yi, Y. A. Kang, J. Stasko, and J. A. Jacko, “Toward a Deeper Understanding of the Role of Interaction in Information Visualization,” IEEE Trans. Visual. Comput. Graphics, vol. 13, no. 6, pp. 1224–1231, 2007, doi: 10.1109/TVCG.2007.70515.</p> <p>[9] E. Dimara and C. Perin, “What is Interaction for Data Visualization?,” IEEE Transactions on Visualization and Computer Graphics, vol. 26, no. 1, pp. 119–129, Jan. 2020, doi: 10.1109/TVCG.2019.2934283.</p> <p>[10] M. Brehmer and T. Munzner, “A Multi-Level Typology of Abstract Visualization Tasks,” IEEE Trans. Visual. Comput. Graphics, vol. 19, no. 12, pp. 2376–2385, 2013, doi: 10.1109/TVCG.2013.124.</p> <p>[11] Y. Shi, T. Gao, X. Jiao, and N. Cao, “Breaking the Fourth Wall of Data Stories Through Interaction,” IEEE Trans. Visual. Comput. Graphics, pp. 1–11, 2022, doi: 10.1109/TVCG.2022.3209409.</p> <p>[12] S. McKenna, N. Henry Riche, B. Lee, J. Boy, and M. Meyer, “Visual Narrative Flow: Exploring Factors Shaping Data Visualization Story Reading Experiences,” Computer Graphics Forum, vol. 36, no. 3, pp. 377–387, 2017, doi: 10.1111/cgf.13195.</p> <p> </p> <h2>Fundings</h2> <p>Project funded by the European Union – NextGenerationEU under the National Recovery and Resilience Plan (NRRP), Investment I.4.1 - Borse PNRR Patrimonio Culturale.</p>
Stable carbon and nitrogen isotope data from Arctic coastscapes associated with coastal invertebrate and fish food webs, 1999-2022
Stable carbon and nitrogen isotope data are commonly used to elucidate food web structure and partition food source importance to consumers. Here, stable isotope data were gathered from across the coastal Arctic to assess how differences in coastal type (i.e., coastscape) and longitudinal region differ across the Arctic. Data were collected between 1999-2022.
Impacts of invasive species on food web energy pathways and quality, St. Lawrence River, 2018-2021.
This dataset contains field measurements collected between 2018 and 2021 from three fluvial lakes in the Upper St. Lawrence River (Canada), including both invaded systems (with dreissenid mussels and round goby) and uninvaded reference sites. Data include georeferenced sampling information (site, lake, latitude, longitude, month, year), water chemistry (total phosphorus, µg/L; conductivity, µS/cm), and habitat descriptors (substrate). Biological records encompass seston, macroinvertebrates, and fish. Fish data comprise species identity, sex, total length (mm), weight (g), relative weight index (Wr), and detailed fatty acid composition expressed as relative proportions (%) and concentrations (µg/mg), including essential LC-PUFAs (EPA, DHA), n-3 and n-6 polyunsaturated fatty acids. Stable isotope data are provided, including carbon (δ13C) and nitrogen (δ15N) ratios, C:N ratios, and isotopic baselines from pelagic (δ13Cpel, δ15Npel) and benthic (δ13Cben, δ15Nben) sources. Derived variables, such as pelagic diet proportion and trophic position, were calculated using the two-source mixing model described by Post (2002) (DOI: https://doi.org/10.1890/0012-9658(2002)083[0703:USITET]2.0.CO;2). These data provide a comprehensive resource for examining food web structure, energy pathways, and the ecological impacts of invasive species in large river ecosystems.
A unified dataset of co-located sewage pollution, periphyton, and benthic macroinvertebrate community and food web structure from Lake Baikal (Siberia)
Sewage released from lakeside development can introduce nutrients and micropollutants that can restructure aquatic ecosystems. Lake Baikal, the world's most ancient, biodiverse, and voluminous lake, has been experiencing localized sewage pollution from lakeside settlements. Increasing filamentous algal abundance suggests benthic communities are responding to this localized pollution. We surveyed 40-km of Lake Baikal's southwestern shoreline 19-23 August 2015 for sewage indicators, including pharmaceuticals, personal care products, and microplastics with co-located periphyton, macroinvertebrate, stable isotope, and fatty acid sampling. Unique identifiers corresponding to sampling locations are retained throughout all data files to facilitate interoperability among the dataset's 150+ variables. The data are structured in a tidy format (a tabular arrangement familiar to limnologists) to encourage future reuse. For Lake Baikal studies, these data can support continued monitoring and research efforts. For global studies of lakes, these data can help characterize sewage prevalence and ecological consequences of anthropogenic disturbance across spatial scales.
Core Research Site Web Quadrat Data for the Net Primary Production Study at the Sevilleta National Wildlife Refuge, New Mexico
This dataset is part of a long-term study at the Sevilleta LTER measuring net primary production (NPP) across four distinct ecosystems: creosote-dominant shrubland (Site C, est. winter 1999), black grama-dominant grassland (Site G, est. winter 1999), blue grama-dominant grassland (Site B, est. winter 2002), and pinon-juniper woodland (Site P, est. winter 2003). Net primary production is a fundamental ecological variable that quantifies rates of carbon consumption and fixation. Estimates of NPP are important in understanding energy flow at a community level as well as spatial and temporal responses to a range of ecological processes. Above-ground net primary production is the change in plant biomass, represented by stems, flowers, fruit and and foliage, over time and incoporates growth as well as loss to death and decomposition. To measure this change the vegetation variables in this dataset, including species composition and the cover and height of individuals, are sampled twice yearly (spring and fall) at permanent 1m x 1m plots within each site. A third sampling at Site C is performed in the winter. The data from these plots is used to build regressions correlating biomass and volume via weights of select harvested species obtained in SEV157, "Net Primary Productivity (NPP) Weight Data." This biomass data is included in SEV182, "Seasonal Biomass and Seasonal and Annual NPP for Core Research Sites." This dataset is designated as NA-US-011 in the Global Index of Vegetation-Plot Databases (GIVD). To aid tracking of the use of databases in this index, please also reference this number when citing this data. The GIVD report for SEV129 can be found in: Biodiversity and Ecology 4 - Vegetation Databases for the 21st Century (2012) by J. Dengler et al.
Core Research Site Web Seasonal Biomass and Seasonal and Annual NPP Data for the Net Primary Production Study at the Sevilleta National Wildlife Refuge, New Mexico
This long-term study at the Sevilleta LTER measures net primary production (NPP) across four distinct ecosystems: creosote-dominant shrubland (Site C, est. winter 1999), black grama-dominant grassland (Site G, est. winter 1999), blue grama-dominant grassland (Site B, est. winter 2002), and pinon-juniper woodland (Site P, est. winter 2003), which is now in its own dataset, SEV278 (Pinon-Juniper (Core Site) Quadrat Data). Net primary production is a fundamental ecological variable that quantifies rates of carbon consumption and fixation. Estimates of NPP are important in understanding energy flow at a community level as well as spatial and temporal responses to a range of ecological processes. While measures of both below- and above-ground biomass are important in estimating total NPP, this study focuses on above-ground net primary production (ANPP). Above-ground net primary production is the change in plant biomass, including loss to death and decomposition, over a given period of time. Volumetric measurements are made using vegetation data from permanent plots collected in SEV129, "Core Research Site Web Quadrat Data" and regressions correlating biomass and volume constructed using seasonal harvest weights from SEV157, "Net Primary Productivity (NPP) Weight Data."
CauseNet: Towards a Causality Graph Extracted from the Web
<p>Causal knowledge is seen as one of the key ingredients to advance artificial intelligence. Yet, few knowledge bases comprise causal knowledge to date, possibly due to significant efforts required for validation. Notwithstanding this challenge, we compile CauseNet, a large-scale knowledge base of <em>claimed </em>causal relations between causal concepts. By extraction from different semi- and unstructured web sources, we collect more than 11 million causal relations with an estimated extraction precision of 83% and construct the first large-scale and open-domain causality graph. We analyze the graph to gain insights about causal beliefs expressed on the web and we demonstrate its benefits in basic causal question answering. Future work may use the graph for causal reasoning, computational argumentation, multi-hop question answering, and more.</p> <p>When using the data, please make sure to refer to it as follows:</p> <pre><code>@inproceedings{heindorf2020causenet, author = {Stefan Heindorf and Yan Scholten and Henning Wachsmuth and Axel-Cyrille Ngonga Ngomo and Martin Potthast}, title = {CauseNet: Towards a Causality Graph Extracted from the Web}, booktitle = {{CIKM}}, pages = {3023--3030}, publisher = {{ACM}}, year = {2020} }</code></pre>
Anonymous Data on Swingers in Germany Harvested on the Web
<p>The data package consists of various files that contain different types of information, mainly focusing on anonymous swingers’ data in various regions:</p> <h2>1. Residents Data (tabular)</h2> <p><strong>Focus</strong>: Demographic and socio-economic data at the county level, focusing on the swinger community. It includes median ages, population<br>densities, and economic factors.<br><strong>Unique Aspects</strong>: Inclusion of demographic details like age groups, employment sectors, and divorce rates, allowing for a deeper socio-economic<br>analysis.<br><strong>Format</strong>: The data are provided in both *.xlsx and *.sav formats, allowing sharing and long-term access to the data.</p> <h2>2. Software</h2> <p>Python scripts used for data conversion and structuring are provided for transparency reasons.</p> <h2>3. Calculation Results Files</h2> <p>Files related to various calculations which had led to the specific design of the data are provided for transparency reasons. They are provided in<br>*.xlsx, *.pdf, and *md format, as is most convenient to adequately reflext the respective content.</p> <p><em><strong>Please refer to the file readme.md for more details.</strong></em></p>
Grapegenomics.com: a web portal with genomic data and analysis tools for wild and cultivated grapevines
<p><a href="https://grapegenomics.com">Grapegenomics.com</a> is a web portal that provides public access to genome references for grapevine cultivars (<em>Vitis vinifera</em> ssp. <em>vinifera</em>), wild grapevines (<em>Vitis vinifera</em> ssp. <em>sylvestris</em>), various wild grape species (<em>Vitis</em> spp. and <em>Muscadinia</em> spp.), and major fungal pathogens affecting grapes.</p> <p>All genomes are accessible through dedicated genome browsers, and published genomes are available for complete <a href="https://www.grapegenomics.com/download.php">download</a>.</p> <p>The site hosts all genomes produced by the laboratory of Dario Cantù in the Department of Viticulture and Enology at the University of California, Davis, along with published genome references generated by others, such as PN40024 and Pinot noir ENTAV115. Instructions for genome submission are provided <a href="https://www.grapegenomics.com/submit.php">here</a>. The portal is maintained by Noé Cochetel (ndcochetel[at]ucdavis.edu). In this version 2.0, all genome browsers utilize <a href="https://jbrowse.org/jb2/">jbrowse 2</a>. <br><br>Link to the website: <a href="https://www.grapegenomics.com">https://www.grapegenomics.com</a> </p>
Density independent prey choice, taxonomy, life history and web characteristics determine the diet and biocontrol potential of spiders (Linyphiidae and Lycosidae) in cereal crops - Dataset
<p>Materials and Methods</p> <p>Fieldwork</p> <p>Money spiders (Araneae: Linyphiidae) and wolf spiders (Araneae: Lycosidae) were the two most common families present in these field surveys, so were prioritised for collection. Spiders were visually located along transects in two adjacent barley fields at Burdons Farm, Wenvoe in South Wales (51°26'24.8"N, 3°16'17.9"W) and collected from occupied webs and the ground, between April and September 2018. Surveys and sampling were conducted five days per week across this period. Each transect was adjacent to a randomly selected tramline and they were distributed across the entire field. The areas searched were 4 m<sup>2</sup> quadrats at least 10 m apart and all observed linyphiids and lycosids were collected in approximately 15-minute searches. The spiders included in this study were taken from 64 locations across 24 days (Supplementary Table 3) along the aforementioned transects. Spiders were individually placed into 1.5 ml microcentrifuge tubes containing 100 % ethanol using an aspirator, regularly changing meshing, at least every five spiders, to limit potential cross-contamination between spiders (spiders were also subsequently washed during transferral to fresh ethanol at the identification and, separately, dissection stages). Linyphiids occupying webs were prioritised for collection, but ground-active linyphiid spiders were also collected. For each spider taken from a web, the height of the web from the ground and its approximate dimensions were recorded, the latter calculated as approximate web area. Spiders were taken to Cardiff University, transferred to fresh ethanol, adults identified to species-level and juveniles to genus, and stored at -80 °C in 100 % ethanol until subsequent DNA extraction. To obtain data on local prey density, 4 m<sup>2</sup> of ground and crop stems were suction sampled using a ‘G-vac’ for 30 seconds at each quadrat from which spiders were collected, with the collected material emptied into a bag, any organisms immediately killed with ethyl-acetate and material frozen for storage before sorting into 70 % ethanol in the lab.</p> <p>All invertebrates were identified to family level due to the restriction of many of the metabarcoding-derived dietary data to this level, and the difficulty associated with finer taxonomic resolution of many taxa. Exceptions included springtails of the superfamily Sminthuroidea (Sminthuridae and Bourletiellidae, which were often indistinguishable following suction sampling and preservation due to the fine features necessary to distinguish them) which were left at super-family, mites (many of which were immature or in poor condition) which were identified to order level and wasps of the superfamily Ichneumonoidea (which were identified no further due to obscurity of wing venation due to damage).</p> <p> </p> <p>Extraction and high-throughput sequencing of spider gut DNA</p> <p>Given their prevalence in field collections, dietary analysis was carried out for the linyphiid genera <em>Erigone</em>, <em>Tenuiphantes</em>, <em>Bathyphantes</em> and <em>Microlinyphia </em>(Araneae: Linyphiidae), and the Lycosidae genus <em>Pardosa</em>. Spiders were transferred to and washed in fresh 100 % ethanol to reduce external contaminants prior to identification via morphological key <sup>1</sup>. Abdomens were removed from spiders and again washed in and transferred to fresh 100 % ethanol. DNA was extracted from the abdomens via Qiagen TissueLyser II and DNeasy Blood & Tissue Kit (Qiagen) as per the manufacturer protocol, but with an extended lysis time of 12 hours to account for the complex and branched gut system in spider abdomens <sup>2</sup>. At least one extraction negative (blank tubes treated identically to samples) was included per 12 spiders (each extraction typically contained 24 spiders, thus two extraction negatives), which was included in subsequent PCR and high-throughput sequencing to detect instances of lab/reagent contamination.</p> <p>For amplification of DNA, two primer pairs were used. BerenF-LuthienR <sup>3</sup> amplified a broad range of invertebrates including spiders, and TelperionF-LaureR, amplified a range of invertebrates but fewer spiders (modified from TelperionF-LaurelinR <sup>3</sup> via one base-pair change from Laurelin; 5’-ggrtawacwgttcawccagt-3’). Primers were labelled with unique 10 bp molecular identifier tags (MID-tags) so that each individual had a unique pairing of forward and reverse tags for identification of each spider post-sequencing. PCR reactions of 25 µl contained 12.5 µl Qiagen PCR Multiplex kit, 0.2 µmol (2.5 µl of 2 µM) of each primer and 5 µl template DNA. Reactions were carried out in the same thermocycler, optimised via temperature gradient, with an initial 15 minutes at 95 °C, 35 cycles of 95 °C for 30 seconds, the primer-specific annealing temperature for 90 seconds and 72 °C for 90 seconds, respectively, followed by a final extension at 72 °C for 10 minutes. BerenF-LuthienR and TelperionF-LaureR used annealing temperatures of 52 °C and 42 °C, respectively.</p> <p>Within each PCR 96-well plate, 12 negative controls (extraction and PCR), 2 blank controls and 2 positive controls were included (i.e. 80 samples per plate), based on Taberlet <em>et al. </em>(2018). Positive controls were mixtures of invertebrate DNA comprised of non-native Asiatic species in four different proportions (Supplementary Table 1) and blanks were empty wells within each plate to identify tag-jumping into unused MID-tag combinations. PCR negative controls were DNase-free water treated identically to DNA samples. A negative control was present for each MID-tag to identify any contamination of primers. All PCR products were visualised in a 2 % agarose gel with SYBRSafe (Thermo Fisher Scientific, Paisley, UK) and placed in categories based on their relative brightness. The concentration of these brightness categories was quantified via Qubit dsDNA High-sensitivity Assay Kits (Thermo Fisher Scientific, Waltham, MA, USA) with at least three representatives of each category per plate. The PCR products were then proportionally pooled according to these concentrations. Each pool was cleaned via SPRIselect beads (Beckman Coulter, Brea, USA), with a left-side size selection using a 1:1 ratio (retaining ~300-1000 bp fragments). The concentration of the pooled DNA was then determined via Qubit dsDNA High-sensitivity Assay Kits and pooled together into one library per primer pair. Library preparation for Illumina sequencing was carried out on the cleaned libraries via NEXTflex Rapid DNA-Seq Kit (Bioo Scientific, Austin, USA) and samples were sequenced on an Illumina MiSeq via a V3 chip with 300-bp paired-end reads (expected capacity ≤25,000,000 reads). Bioinformatic analysis followed (Drake et al., 2021; Supplementary Information 1).</p> <p> </p> <p>Statistical analysis</p> <p>All analyses were conducted in R v4.0.0 <sup>6</sup>. Initial multivariate analyses used binary data (i.e., presence/absence) given the various problems inherent to quantifying metabarcoding data <sup>7,8</sup>. Prey species that occurred only once across all of the dietary samples were removed before further analyses to prevent outliers skewing the results, which is particularly problematic for non-metric multidimensional scaling. Spider diets were compared between variables using multivariate generalized linear models (MGLMs) via ‘manyglm’ in the ‘mvabund’ package <sup>9</sup> with a binomial error family and Monte Carlo resampling. Model independent variables included spider genus, spider life stage (juvenile or adult, the latter defined by fully developed genitalia), spider sex and all two-way interactions between these variables. Pairwise two-way interactions were also included between the aforementioned variables and Julian day to account for how seasonality may affect these relationships.</p> <p>Coarse dietary differences were visualised by non-metric multidimensional scaling (NMDS) via metaMDS in the ‘vegan’ package <sup>10</sup> with Jaccard distance in two dimensions and 999 tries. For NMDS, outliers (usually samples containing rare taxa) were identified by plotting and subsequently removed to facilitate separation of samples and achieve minimum stress. For visualisation of the effect of categorical variables against the dietary NMDS, spider plots were created using ‘ordispider’ with ‘ggplot’ and the ‘RColorBrewer’ ‘Accent’ colour palette <sup>11</sup>. Spider diet was compared against web characteristics for spiders for which both data were available using the MGLM process outlined above, but with starting models containing web height, web area, an interaction between the two, and pairwise interactions between genus, life stage and sex with the two web variables. This model used the same binomial error family as above, but with a ‘cloglog’ link function. For visualisation of the effect of continuous variables against the NMDS, surf plots were created with scaled coloured contours using the function “ordisurf” of the “ggplot” package in R.</p> <p>All prey taxa were classified as agricultural pests, natural enemies or excluded from subsequent analyses of intraguild predation and biocontrol (Supplementary Table 2). Intraguild predation and biocontrol variables were created by counting the number of natural enemy taxa, and, separately, of agriculturally relevant “pest” taxa (taxa containing species that commonly adversely affect agricultural productivity; Supplementary Table 2) in each spider’s diet. These resultant count data (effectively the diversity of pests and natural enemies predated by each individual spider) were separately analysed against spider genus, life stage and sex via GLM. “Site” (denoting the 4 m<sup>2</sup> area from which spiders were collected within fields) was initially included as a random effect in generalized linear mixed-models, but no significant effect was observed when comparing this model against a standard GLM via a likelihood ratio test of nested models using the ‘lrtest’ command in the ‘lmtest’ package <sup>12</sup>. Standard GLMs were thus used to avoid issues relating to singularity in the mixed models. The assumptions for the resultant Poisson error family GLMs were tested using the “testResiduals” function of the ‘DHARMa’ package <sup>13</sup>. Intraguild predation and biocontrol differences between significant terms were visualised using violin plots with the quartiles, median and 95 % upper limit annotated using the ‘geom_violin’ function in ‘ggplot2’.</p> <p><em>In situ</em> spider prey choice was analysed using network-based null models in the ‘econullnetr’ package <sup>14</sup> with the ‘generate_null_net’ command, visually represented with the ‘plot_preferences’ command. Binary dietary data were used alongside suction sample count data to represent prey availability. These suction sample data, as described above, were collected at the same sites as the spiders three days after spider collection. Prior to the taxonomic prey choice analysis, an hemipteran identified no further than order level through dietary analysis was removed due to the inability to pair it to any present prey taxa with certainty. Standardised effect sizes (SES) were extracted for all comparisons for each individual spider and compared between genera, life stages and sexes using permutational multivariate analysis of variance (PerMANOVA) using the ‘adonis’ function of the ’vegan’ package with 9999 permutations and a Euclidean distance matrix to determine overall differences in prey choice.</p> <p> </p> <p>References</p> <p>1. Roberts, M. J. <em>The Spiders of Great Britain and Ireland (Compact Edition)</em>. (Harley Books, 1993).</p> <p>2. Krehenwinkel, H., Kennedy, S., Pekár, S. & Gillespie, R. G. A cost-efficient and simple protocol to enrich prey DNA from extractions of predatory arthropods for large-scale gut content analysis by Illumina sequencing. <em>Methods Ecol. Evol.</em> <strong>8</strong>, 126–134 (2017).</p> <p>3. Cuff, J. P. <em>et al.</em> Money spider dietary choice in pre- and post-harvest cereal crops using metabarcoding. <em>Ecol. Entomol.</em> <strong>46</strong>, 249–261 (2021).</p> <p>4. Taberlet, P., Bonin, A., Zinger, L. & Coissac, E. <em>Environmental DNA</em>. (Oxford University Press, 2018).</p> <p>5. Drake, L. E. <em>et al.</em> An assessment of minimum sequence copy thresholds for identifying and reducing the prevalence of artefacts in dietary metabarcoding data. <em>Methods Ecol. Evol.</em> <strong>in press</strong>, (2021).</p> <p>6. R Core Team. R: A language and environment for statistical computing. (2020).</p> <p>7. Deagle, B. E., Thomas, A. C., Shaffer, A. K. & Trites, A. W. Quantifying sequence proportions in a DNA-based diet study using Ion Torrent amplicon sequencing: which counts count? <em>Mol. Ecol. Resour.</em> <strong>13</strong>, 620–633 (2013).</p> <p>8. Deagle, B. E. <em>et al.</em> Counting with DNA in metabarcoding studies: How should we convert sequence reads to dietary data? <em>Mol. Ecol.</em> <strong>28</strong>, 391–406 (2019).</p> <p>9. Wang, Y., Naumann, U., Wright, S. T. & Warton, D. I. mvabund – an R package for model-based analysis of multivariate abundance data. <em>Methods Ecol. Evol.</em> <strong>3</strong>, 471–474 (2012).</p> <p>10. Oksanen, J. <em>et al.</em> vegan: Community Ecology Package. (2016).</p> <p>11. Neuwirth, E. RColorBrewer: ColorBrewer palettes. (2014).</p> <p>12. Zeileis, A. & Hothorn, T. Diagnostic checking in regression relationships. <em>R News</em> <strong>2</strong>, 7–10 (2002).</p> <p>13. Hartig, F. DHARMa: residual diagnostics for hierarchical (multi-level/mixed) regression models. (2020).</p> <p>14. Vaughan, I. P. <em>et al.</em> econullnetr: an r package using null models to analyse the structure of ecological networks and identify resource selection. <em>Methods Ecol. Evol.</em> <strong>9</strong>, 728–733 (2018).</p>
Annual Article Processing Charges (APCs) and number of gold and hybrid open access articles in Web of Science indexed journals published by Elsevier, Sage, Springer-Nature, Taylor & Francis and Wiley 2015-2018
<p><strong>Dataset of annual Article Processing Charges (APCs) for 6,252 journals from 2015 to 2018. </strong>The dataset contains annual APCs for journals indexed in the Web of Science (WoS) and published by the oligopoly of academic publishers (Elsevier, Sage, Springer-Nature, Taylor & Francis, Wiley). It also includes an estimate of the total APCs paid by the academic community based on the number of gold and hybrid articles published between 2015 and 2018. The dataset was created using publication data from WoS, OA status from Unpaywall and annual APC prices from open datasets (<a href="https://doi.org/10.5281/ZENODO.3841568">Matthias, 2020</a>; <a href="https://doi.org/10.5683/SP2/84PNSG">Morrison, 2021</a>) and historical fees retrieved via the Internet Archive Wayback Machine. </p> <p>Detailed methods and findings are reported in the following journal article</p> <p>Butler, L.-A., Matthias, L., Simard, M.-A., Mongeon, P., & Haustein, S. (2023). The Oligopoly's Shift to Open Access. How the Big Five Academic Publishers Profit from Article Processing Charges. <em>Quantitative Science Studies</em>. Preprint: <a href="https://doi.org/10.5281/zenodo.8322555">https://doi.org/10.5281/zenodo.8322555</a></p> <p><strong>Description of included files (v1):</strong></p> <p><em>APCs.csv: </em>contains the annual APCs for gold and hybrid OA journals indexed in Web of Science published by the oligopoly of academic publishers (Elsevier, Sage, Springer-Nature, Taylor & Francis, Wiley) between 2015 and 2018 including the total estimate of APCs paid per journal per year. It contains APC data for 18,846 journal-year-OA status combinations.</p> <p><em>countries.csv</em>: contains the fractionalized number of annual gold and hybrid OA articles by oligopoly publishers between 2015 and 2018 and the total estimate of fractionalized APCs paid per country per journal per year.</p> <p><em>oecd.csv</em>: contains the fractionalized number of annual gold and hybrid OA articles by oligopoly publishers between 2015 and 2018 and the total estimate of fractionalized APCs per discipline per journal per year.</p> <p><em>ReadMe.csv</em>: contains a description of the variables used in <em>APCs.csv</em>, <em>countries.csv</em> and <em>oecd.csv</em>.</p> <p> </p>
Web Experience in Mobile Networks: Lessons from Two Million Page Visits
<p>Measuring and characterizing web page performance is a challenging task.</p> <p>When it comes to the mobile world, the highly varying technology characteristics coupled with the opaque network configuration make it even more difficult.</p> <p>Aiming at reproducibility, we present a large scale measurements study of web page performance collected in eleven commercial mobile networks spanning four countries.</p> <p>We build a dataset of nearly two million web browsing sessions to we shed light on the impact of different web protocols, browsers, and mobile technologies on the web performance.</p> <p>We find that the impact of mobile broadband access is sizeable.</p> <p>For example, the median page load time using mobile broadband increases by a third compared to wired access.</p> <p>Mobility clearly stresses the system, with handover causing the most evident performance penalties.</p> <p>Contrariwise, our measurements show that the adoption of HTTP/2 and QUIC has practically negligible impact.</p> <p>Our work highlights the importance of large-scale measurements.</p> <p>Even with our controlled setup, the complexity of the mobile web ecosystem is challenging to untangle.</p> <p>For this, we are releasing the dataset as open data for validation and further research.</p> <p>We also release together with the datasets we collected the scripts we use to produce the analysis we present in the paper. Please use plot_all.sh script to generate the plots in the paper, using the separate scripts from the "scripts" archive. </p> <p>Should you use any of these resources, please also make an attribution using the following reference (provided here in bibtex format):</p> <pre>@inproceedings{rajiullah2019web, title={{Web Experience in Mobile Networks: Lessons from Two Million Page Visits}}, author={Rajiullah, Mohammad and Lutu, Andra and Khatouni, Ali Safari and Fida, Mah-Rukh and Mellia, Marco and Brunstrom, Anna and Alay, Ozgu and Alfredsson, Stefan and Mancuso, Vincenzo}, booktitle={The World Wide Web Conference}, pages={1532--1543}, year={2019}, organization={ACM}, address = {San Francisco, CA, USA}, keywords = {Web Experience, HTTP2, QUIC, TCP, Mobile Broadband, Measurements} }</pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.