Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,359

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,359 results for “Online”

Learn how ShareScore rates datasets ↗
zenodo44/100

The Online Conversation Threads Repository (Slashdot, Barrapunto, Wikipedia talk)

<p>This repository contains datasets with online conversation threads collected and analyzed by different researchers. Currently, you can find datsets from different news aggregators (Slashdot, Barrapunto) and the English Wikipedia talk pages.</p> <p>- Slashdot conversations (Aug 2005 - Aug 2006) Online conversations generated at Slashdot during a year. Posts and comments published between August 26th, 2005 and August 31th, 2006. For each discussion thread: sub-domains, title, topics and hierarchical relations between comments. For each comment: user, date, score and textual content. This dataset is different from the Slashdot Zoo social network (it is not a signed network of users) contained in the SNAP repository and represents the full version of the dataset used in the CAW 2.0 - Content Analysis for the WEB 2.0 workshop for the WWW 2009 conference that can be found in several repositories such as Konect Barrapunto conversations (Jan 2005 - Dec 2008)</p> <p>- Online conversations generated at Barrapunto (Spanish clone of Slashdot) during three years. For each discussion thread: sub-domains, title, topics and hierarchical relations between comments. For each comment: user, date, score and textual content Wikipedia (2001 - Mar 2010)</p> <p>- Data from articles discussions (talk) pages of the English Wikipedia as of March 2010. It contains comments on about 870,000 articles (i.e. all articles which had a corresponding talk page with at least one comment), in total about 9.4 million comments. The oldest comments date back to as early as 2001.</p> <p> </p>

opencc-by-4.0Apr 2008View details →
zenodo44/100

World Flora Online Plant List June 2025

<p>The consensus taxonomy of plants used as the backbone for the <a href="https://www.worldfloraonline.org/">World Flora Online</a> (WFO) portal, and issued as editions of the <a href="https://wfoplantlist.org/">WFO Plant List</a>.</p> <p>New versions of this checklist are released every six months in June and December: this is release 2025-06.</p> <p>The history of data development for the WFO taxonomic backbone is given on the WFO Plant List <a href="https://wfoplantlist.org/background">background page</a>. Taxonomic names are incorporated into WFO from nomenclators <a href="https://www.ipni.org/">International Plant Name Index</a> (IPNI) for vascular plants, and <a href="https://www.tropicos.org/home">Tropicos</a> for bryophytes. Taxonomic and nomenclatural updates are incorporated from the WFO's <a href="https://about.worldfloraonline.org/tens">Taxonomic Expert Networks</a> (TENs) and the <a href="https://powo.science.kew.org/about-wcvp">World Checklist of Vascular Plants</a> (WCVP), facilitated by the Royal Botanic Gardens, Kew.</p> <p>This data repository includes the following files:</p> <ul> <li><strong>wfo_plantlist_2025-06.zip</strong> The Catalogue of Life Data Package of the WFO Plant List. This is the most expressive standards based form of the list.</li> <li><strong>plant_list_2025-06.json.gz</strong> JSON formatted version of the WFO Plant List. This has been designed for direct import into a schemaless instance of a SOLR index and is used to drive the WFO Plant List API (<a href="https://list.worldfloraonline.org">https://list.worldfloraonline.org</a>) which in turn drives the WFO Plant List in the portal. This is recommended if you want a local, read only version of the list rather than use the API.</li> <li><strong>plant_list_2025-06.sql.gz</strong> This is the complete production database (minus API keys) as a MySQL backup file. It can be restored directly to a MySQL 8.0 or later instance if you require the list in SQL format.</li> <li><strong>ipni_to_wfo.csv.gz</strong> A file mapping all the IPNI IDs we track to their associated WFO IDs.</li> <li><strong>families_dwc.tar.gz</strong> Individual Darwin Core Archive files for each of 733 recognized families. If you want a single family in DwC but can't load the whole list download and expand this file. Family and genus files are also available for download through the portal. These files exclude deprecated names.</li> <li><strong>_DwC_backbone_R.zip</strong> A single Darwin Core Archive file containing non deprecated names and taxa for use in the existing R package.</li> <li><strong>_uber.zip</strong> A single Darwin Core Archive file containing all names and taxa even those that are deprecated along with some extra columns</li> </ul>

opencc-zeroJun 2024View details →
zenodo44/100

Productos multivitamínicos en stock de la tienda online HSN con sus descuentos, características y valoraciones

<p>Los datos extraídos mediante web scraping de la web de HSN nos proporcionan información sobre los productos multivitamínicos que ofrece esta tienda online. En el dataset se recogen los diferentes productos que ofrecen y en sus diferentes tamaños, siempre y cuando alguno de los tamaños disponibles para el producto esté en stock.&nbsp;</p><ul><li><strong>Código</strong>: Referencia del producto de HSN, diferentes formatos del mismo producto tienen el mismo código al pertenecer al mismo producto.</li><li><strong>Nombre</strong>: Nombre del producto</li><li><strong>Clasificación</strong>: Categoría a la que pertenece el producto</li><li><strong>Grupo</strong>: Subcategoría a la que pertenece el producto</li><li><strong>Dosificación</strong>: Formato o tamaño en el que se presenta el producto</li><li><strong>Precio</strong>: Precio original del producto antes del descuento</li><li><strong>Precio_descuento</strong>: Precio final del producto tras realizar el descuento</li><li><strong>%_Descuento</strong>: Porcentaje de descuento sobre el precio original</li><li><strong>Rating</strong>: Valoración sobre 5 del producto según los consumidores</li><li><strong>Numero_valoraciones</strong>: Número de valoraciones realizadas sobre el producto</li><li><strong>Serie</strong>: Serie a la que pertenece el producto</li><li><strong>Descripción</strong>: Descripción del producto</li><li><strong>Edulcorantes</strong>: Recoge si el producto incluye o no edulcorantes</li><li><strong>Veganos</strong>: Recoge si el producto es apto para veganos</li><li><strong>Vegetarianos</strong>: Recoge si el producto es apto para vegetarianos</li><li><strong>OMG</strong>: Recoge si el producto contiene organismos modificados genéticamente</li><li><strong>Gluten</strong>: Recoge si el producto contiene gluten</li><li><strong>Lactosa</strong>: Recoge si el producto contiene lactosa</li><li><strong>Lácteos</strong>: Recoge si el producto contiene lácteos</li><li><strong>Soja</strong>: Recoge si el producto contiene soja</li><li><strong>Huevo</strong>: Recoge si el producto contiene huevo</li><li><strong>Conservantes</strong>: Recoge si el producto contiene conservantes</li><li><strong>Colorantes</strong>: Recoge si el producto contiene colorantes artificiales</li><li><strong>Frutos_cascara</strong>: Recoge si el producto contiene frutos de cáscara</li><li><strong>Pescado</strong>: Recoge si el producto contiene pescado</li></ul><p>En cuanto al periodo de tiempo al que pertenecen los datos (02-11-2023) hay que destacar que la página está constantemente publicando nuevas ofertas, por lo que se podría considerar que los datos son válidos en ese momento, y que posteriormente a la recogida de datos la página puede publicar nuevas ofertas.</p>

opencc-by-sa-4.0Nov 2023View details →
zenodo44/100

Phlorest phylogeny derived from Greenhill 2015 'TransNewGuinea.org: An Online Database of New Guinea Languages'

<p>Cite the source of the dataset as:</p> <blockquote> <p>Greenhill, S. J. (2015). TransNewGuinea.org: An Online Database of New Guinea Languages. PLOS ONE, 10(10), e0141563. doi:10.1371/journal.pone.0141563</p> </blockquote>

opencc-by-4.0Aug 2023View details →
zenodo44/100

Country Compendium of the Global Register of Introduced and Invasive Species: Standardization to Records in World Flora Online or the World Checklist of Vascular Plants

<p>The <strong>Country Compendium of the Global Register of Introduced and Invasive Species (GRIIS)</strong> is a collation of data across 196 individual country checklists of alien species, along with a designation of those species associated with evidence of impact at a country level. This compendium is available via <a href="https://zenodo.org/records/6348164">Zenodo</a> and was described by Pagad et al. <a href="https://www.nature.com/articles/s41597-022-01514-z">2022</a>:</p><ul><li>Shyama Pagad, Stewart Bisset, &amp; Melodie A. McGeoch. (2022). Country Compendium of the Global Register of Introduced and Invasive Species. Dataset. (V1_0) [Data set]. Zenodo. <a href="https://doi.org/10.5281/zenodo.6348164">https://doi.org/10.5281/zenodo.6348164</a></li><li>Pagad, S., Bisset, S., Genovesi, P. <i>et al.</i> Country Compendium of the Global Register of Introduced and Invasive Species. <i>Sci Data</i> <strong>9</strong>, 391 (2022). <a href="https://doi.org/10.1038/s41597-022-01514-z">https://doi.org/10.1038/s41597-022-01514-z</a></li></ul><p>&nbsp;</p><p>Here I provide direct and fuzzy matches for species listed for the Plantae Kingdom in GRIIS with accepted plant names in <strong>World Flora Online</strong> (<a href="https://www.worldfloraonline.org/downloadData">version 2023.03</a>; Borsch et al. <a href="https://doi.org/10.1002/tax.12373">2020</a>) or the <strong>World Checklist of Vascular Plants</strong> (<a href="https://doi.org/10.34885/nswv-8994">version 10</a>; Govaerts et al. <a href="https://www.nature.com/articles/s41597-021-00997-6">2021</a>). Matching was done in <i>R</i> through the <a href="https://cran.r-project.org/package=WorldFlora">WorldFlora</a> package (Kindt <a href="https://bsapubs.onlinelibrary.wiley.com/doi/full/10.1002/aps3.11388">2020</a>). The taxonomic standardization process was similar to the one completed <a href="https://www.worldagroforestry.org/output/agroforestry-species-switchboard-30">during the preparation of the third major release</a> of the <a href="https://apps.worldagroforestry.org/products/switchboard">Agroforestry Species Switchboard</a> and when preparing the <strong>GlobalUsefulNativeTrees database</strong> (GlobUNT; <a href="https://worldagroforestry.org/output/globalusefulnativetrees">https://worldagroforestry.org/output/globalusefulnativetrees</a>) .</p><p>Where a matching species was found in GlobUNT, the species name in the GlobUNT database has been shown. GlobUNT has been described in the following publication: Kindt et al. (<a href="https://www.nature.com/articles/s41598-023-39552-1">2023</a>) <strong>GlobalUsefulNativeTrees, a database of 14,014 tree species, supports synergies between biodiversity recovery and local livelihoods in restoration</strong>. <i>Sci Rep</i> <strong>13</strong>, 12640. <a href="https://doi.org/10.1038/s41598-023-39552-1">https://doi.org/10.1038/s41598-023-39552-1</a>.</p><p>The developments of this dataset and GlobUNT were supported by the Darwin Initiative to project DAREX001 of <a href="https://www.darwininitiative.org.uk/project/DAREX001/"><i>Developing a Global Biodiversity Standard certification for tree-planting and restoration</i></a> and by Norway's International Climate and Forest Initiative through the Royal Norwegian Embassy in Ethiopia to the <a href="https://www.worldagroforestry.org/project/provision-adequate-tree-seed-portfolio-ethiopia"><i>Provision of Adequate Tree Seed Portfolio</i></a> project in Ethiopia.&nbsp;</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Toxic Content Detection in online social networks: a new dataset from Brazilian Reddit Communities

<p>This is new dataset of 2,500 manually annotated examples of comments extracted from the top 10 largest Brazilian subreddits on Reddit. The dataset has been annotated by crowd-sourcing efforts with contributions from the departments of computer science (DCC) and the linguistic group @ UFMG. As part of our contribution to the toxicity automatic detection and moderation of online social networks, we're making the dataset public for research.</p> <h3>Dataset</h3> <p>The dataset contains 2,500 manually annotated comments from the most popular brazilian communities on Reddit. The data sampling proccess was a stratified sampling by the number of generated publications by subreddit and the month of publication. The list of communities collected is presented below. The collected data period ranges from January 2022 to December 2022.</p> <p>&nbsp;</p> <table> <tbody> <tr> <td><strong>Subreddit</strong></td> <td><strong>Posts</strong></td> <td><strong>Comments</strong></td> </tr> <tr> <td>r/brasil</td> <td>110,829&nbsp;</td> <td>2,136,866</td> </tr> <tr> <td>r/desabafos</td> <td>115,876</td> <td>1,211,643</td> </tr> <tr> <td>r/futebol</td> <td>35,826</td> <td>1,214,412</td> </tr> <tr> <td>r/saopaulo</td> <td>7,308</td> <td>81,969</td> </tr> <tr> <td>r/eu_nvr</td> <td>12,631</td> <td>188,620</td> </tr> <tr> <td>r/botecodoreddit</td> <td>7,059</td> <td>57,298</td> </tr> <tr> <td>r/conversas</td> <td>21,967</td> <td>326,061</td> </tr> <tr> <td>r/investimentos</td> <td>9,756</td> <td>141,823</td> </tr> <tr> <td>r/tiodopave</td> <td>2,371</td> <td>11,584</td> </tr> <tr> <td>r/brasilivre</td> <td>67,301</td> <td>1,219265</td> </tr> <tr> <td>Total</td> <td>390,924</td> <td>6,589,541</td> </tr> </tbody> </table> <p>&nbsp;</p> <h3><strong>Annotation proccess</strong></h3> <p>The annotators were divided into groups of raters and each group was assigned a batch of comments to label. The raters were then asked to label a comment as <strong>Toxic</strong>, <strong>Non-toxic</strong>, <strong>I do not know</strong> and <strong>Missing info</strong>. During the annotation process, the raters were encouraged to assign one of the uncertain labels when they're not sure about the toxicity of a comment or the context is missing.&nbsp;</p> <h3>Available data</h3> <p>The dataset is available as csv file and the label was assigned as a majority vote among the raters. The available data are the original collected comment id and body. The label was created from the original classification from the annotators. No data processing has been done on this version of the dataset. The overall schema of the dataset if presented below.</p> <p>- <strong>id</strong>: The unique identifier of the comment on the Reddit platform<br>- <strong>body</strong>: The original comment text publication<br>- <strong>is_toxic</strong>: The final label of a given comment. The label is <strong>0</strong> for non-toxic comments, <strong>1</strong> for toxic comments and <strong>-1</strong> for comments where the raters disagreed about the toxicity.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Trees of India Version 1: Standardization to Records in World Flora Online and the World Checklist of Vascular Plants, with matches in GlobalTreeSearch and GlobalUsefulNativeTrees

<p>The <strong>Trees of India (ToI, Version-I)</strong> includes data on 3708 tree species distributed across 35 states/union territories of India. The database is based on systematic review of 313 literature sources published from 1872-2022.This compendium is available via <a href="https://figshare.com/articles/dataset/ToI_Ver_-I_Trees_of_India_Version-I/23226281">Figshare</a> and was described by Mugal et al. <a href="https://link.springer.com/article/10.1007/s10531-023-02659-y">2023</a>:</p> <ul> <li>Khuroo, Anzar Ahmad; Mugal, Muzamil Ahmad; Wani, Sajad Ahmad (2023). ToI, Ver.-I : Trees of India, Version-I. figshare. Dataset. <a href="https://doi.org/10.6084/m9.figshare.23226281.v1">https://doi.org/10.6084/m9.figshare.23226281.v1</a></li> <li>Mugal, M.A., Wani, S.A., Dar, F.A. <em>et al.</em> Bridging global knowledge gaps in biodiversity databases: a comprehensive data synthesis on tree diversity of India. <em>Biodivers Conserv</em> <strong>32</strong>, 3089&ndash;3107 (2023). <a href="https://doi.org/10.1007/s10531-023-02659-y">https://doi.org/10.1007/s10531-023-02659-y</a></li> </ul> <p>&nbsp;</p> <p>Here I provide direct and fuzzy matches for taxa listed with accepted plant names in <strong>World Flora Online</strong> (<a href="https://www.worldfloraonline.org/downloadData">version 2023.03</a>; Borsch et al. <a href="https://doi.org/10.1002/tax.12373">2020</a>) and the <strong>World Checklist of Vascular Plants</strong> (WCVP <a href="https://doi.org/10.34885/nswv-8994">version 10</a>; Govaerts et al. <a href="https://www.nature.com/articles/s41597-021-00997-6">2021</a>). Matching was done in <em>R</em> through the <a href="https://cran.r-project.org/package=WorldFlora">WorldFlora</a> package (Kindt <a href="https://bsapubs.onlinelibrary.wiley.com/doi/full/10.1002/aps3.11388">2020</a>). The taxonomic standardization process was similar to the one completed <a href="https://www.worldagroforestry.org/output/agroforestry-species-switchboard-30">during the preparation of the third major release</a> of the <a href="https://apps.worldagroforestry.org/products/switchboard">Agroforestry Species Switchboard</a> and when preparing the <strong>GlobalUsefulNativeTrees database</strong> (GlobUNT; <a href="https://worldagroforestry.org/output/globalusefulnativetrees">https://worldagroforestry.org/output/globalusefulnativetrees</a>).</p> <p>After matching species with the WCVP, information was compiled on the <strong>native distribution</strong> documented in the WCVP for level-3 units of the <a href="https://github.com/tdwg/wgsrpd">World Geographical Scheme for Recording Plant Distributions</a> that correspond to India, including India (IND), Assam (ASS), West Himalaya (WHM), East Himalaya (EHM), Laccadive Is. (LDV), Andaman Is. (AND) and Nicobar Is. (NCB). Also included after matching with the WCVP is information on the geographic area, lifeform and main biome. Similar information is available when searching for species from <a href="https://powo.science.kew.org/">Plants of the World Online</a>.</p> <p>Where a matching species was found in <strong>GlobalTreeSearch</strong> (Beech et al. <a href="https://www.tandfonline.com/doi/full/10.1080/10549811.2017.1310049">2017</a>; <a href="https://tools.bgci.org/global_tree_search.php">https://tools.bgci.org/global_tree_search.php</a>; accessed on 28th June 2023) filtered for India, the species name in GlobalTreeSearch is shown. Note that GlobalTreeSearch documents the <strong>native country distribution</strong> of tree species.</p> <p>Where a matching species was found in the <strong>GlobalUsefulNativeTrees</strong> database (GlobUNT, version 2023.11) filtered for India, the species name in the GlobUNT database is shown. GlobUNT has been described in the following publication: Kindt et al. (<a href="https://www.nature.com/articles/s41598-023-39552-1">2023</a>) <strong>GlobalUsefulNativeTrees, a database of 14,014 tree species, supports synergies between biodiversity recovery and local livelihoods in restoration</strong>. <em>Sci Rep</em> <strong>13</strong>, 12640. <a href="https://doi.org/10.1038/s41598-023-39552-1">https://doi.org/10.1038/s41598-023-39552-1</a>.</p> <p>See the metadata for information on versions.</p> <p>&nbsp;</p> <ul> <li>Borsch, T., Berendsohn, W., Dalcin, E., Delmas, M., Demissew, S., Elliott, A., Fritsch, P., Fuchs, A., Geltman, D., G&uuml;ner, A., Haevermans, T., Knapp, S., le Roux, M.M., Loizeau, P.-A., Miller, C., Miller, J., Miller, J.T., Palese, R., Paton, A., Parnell, J., Pendry, C., Qin, H.-N., Sosa, V., Sosef, M., von Raab-Straube, E., Ranwashe, F., Raz, L., Salimov, R., Smets, E., Thiers, B., Thomas, W., Tulig, M., Ulate, W., Ung, V., Watson, M., Jackson, P.W. and Zamora, N. (2020), World Flora Online: Placing taxonomists at the heart of a definitive and comprehensive global resource on the world's plants. TAXON, 69: 1311-1341. <a href="https://doi.org/10.1002/tax.12373">https://doi.org/10.1002/tax.12373</a></li> <li>Govaerts, R., Nic Lughadha, E., Black, N. <em>et al.</em> The World Checklist of Vascular Plants, a continuously updated resource for exploring global plant diversity. <em>Sci Data</em> <strong>8</strong>, 215 (2021). <a href="https://doi.org/10.1038/s41597-021-00997-6">https://doi.org/10.1038/s41597-021-00997-6</a></li> <li>E.&nbsp;Beech,&nbsp;M.Rivers,&nbsp;S.&nbsp;Oldfield &amp;&nbsp;P. P.&nbsp;Smith (2017)GlobalTreeSearch: The first complete global database of tree species and country distributions, Journal of Sustainable Forestry, 36:5, 454-489, DOI: <a href="https://doi.org/10.1080/10549811.2017.1310049">10.1080/10549811.2017.1310049</a></li> <li>Kindt, R. 2020. WorldFlora: An R package for exact and fuzzy matching of plant names against the World Flora Online taxonomic backbone data. <em>Applications in Plant Sciences</em> 8(9): e11388. <a href="https://doi.org/10.1002/aps3.11388">https://doi.org/10.1002/aps3.11388</a></li> </ul> <p>&nbsp;</p> <p>The developments of this dataset and GlobUNT were supported by the Darwin Initiative to project DAREX001 of <a href="https://www.darwininitiative.org.uk/project/DAREX001/"><em>Developing a Global Biodiversity Standard certification for tree-planting and restoration</em></a>.</p>

opencc-by-4.0Nov 2023View details →
zenodo44/100

Questionnaire survey among members of the Czech Pirate Party regarding the use of the online voting system Helios

<p>LimeSurvey application, where only Pirate Party members had access to the survey via a unique URL sent in an e-mail invitation that allowed members to fill out the questionnaire once. The poll ran from 22 February to 14 March 2024. 213 members out of a total of 1,179 party members completed the 19-question poll in full, a response rate of 18.6%. 16 research questions were Yes or No answers, 3 questions were socio-demographic questions focusing on gender, age and educational attainment. 47 women, 146 men and 20 respondents did not classify themselves as male or female. The age group 18-30 years included 27 respondents, 31-40 years included 93 respondents, 41-50 years included 56 respondents, 51-60 years included 24 respondents, 61-70 years included 11 respondents, 71 years and above included 2 respondents. In terms of highest completed education of the respondents, 60 have high school degree, 27 have bachelor's degree, 97 have master's degree, 13 have doctoral degree and 16 have other degree.</p>

opencc-by-4.0Mar 2024View details →
zenodo44/100

Online Knowledge Production in Polarized Political Memes: The Case of Critical Race Theory (Dataset)

<p>This is the supplementary dataset to the article entitled "Online Knowledge Production in Polarized Political Memes: The Case of Critical Race Theory." This study, completed by Alyvia Walters, Tawfiq Ammari, Kiran Garimella, and Shagun Jhaver, was accepted for publication in <em>New Media &amp; Society&nbsp;</em>in 2024.</p> <p>Description of files:</p> <p>Memes Codebook.dox - Codebook used for qualitative coding of memes.</p> <p>Memes Project.qdpx - NVivo coding project, exported.</p> <p>all_posts.jsonl, clusters.zip, images.zip, &amp; image_data.csv - the complete collection of data and images used for this project. Please see publication in&nbsp;<em>New Media &amp; Society&nbsp;</em>for more information on the use and collection of these items.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Challenges to freedom of speech and journalists in Ukraine in times of war – Non-representative online expert survey of Ukrainian journalists (January 2023)

The expert survey of journalists was conducted from 18 to 27 January 2023 using a self-completion questionnaire in Google Forms. The survey was conducted by the Ilko Kucheriv Democratic Initiatives Foundation on the request of the Human Rights Centre ZMINA with the support of Freedom House Ukraine. A total of 132 people participated in the survey. The respondents were selected using the method of voluntary selection and snowballing to the point of saturation. The sample represents only the opinion of the respondents, but it also allows us to talk about certain trends and common assessments of certain phenomena and processes in the journalistic field. The survey includes questions about freedom of speech and self-censorship in the media environment during the Russian-Ukrainian war. The data collection contains original survey data. The Excel file (.xlsx) is the original file with the respondents' answers in Ukrainian, provided by the Ilko Kucheriv Democratic Initiatives Foundation. The documentation includes the questions and answer options of the original questionnaire in Ukrainian and English. Additionally, the data collection contains the "Summary" file, which is an analytical report prepared by the Ilko Kucheriv Democratic Initiatives Foundation and the Human Rights Centre ZMINA. The report uses data from an expert survey of journalists in 2019 and 2023, and the results of focus groups in 2022.

openodc-byDec 2024View details →
zenodo44/100

Evaluation datasets and results of the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing"

<p>Event logs, process models, and results corresponding to the paper "Efficient Online Computation of Business Process State From Trace Prefixes via N-Gram Indexing".</p> <p><em><strong>Inputs</strong></em>: preprocessed event logs and discovered process models (and their characteristics) used in the evaluation.</p> <ul> <li><em><strong>Real-life</strong></em>: preprocessed event logs (<em>xes</em> and <em>csv</em>) corresponding to the real-life processes used in the evaluation. Process models (<em>pnml</em>) discovered with the Inductive Miner infrequent for thresholds of 10%, 20%, and 50%. Characteristics (<em>txt</em>) of the event logs and process models. Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>).</li> <li><em><strong>Synthetic</strong></em>: simulated&nbsp;event logs (<em>csv</em>) corresponding to the synthetic processes used in the evaluation. Designed process models (<em>bpmn</em> and&nbsp;<em>pnml</em>). Ongoing cases result from splitting each case in the preprocessed event logs (under folder <em>split</em>). Ongoing cases with injected noise as described in the publication (under folders <em>noise_1</em>, <em>noise_2</em>, and <em>noise_3</em>).</li> </ul>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Social Network Online Activity of 100+ Users Over Two Years

<p>This dataset contains a precise (error margin is within&nbsp;5 seconds) activity log of 138 users recorded over a period of approximately two years. It includes users&#39;&nbsp;log in/log off timestamps as well as a device id which was used during the session. An activity heat map is also provided which can be used to determine the online time (in seconds) in a given hour for a given user. The dataset is completely anonymized and is not linked to real peoples&#39;&nbsp;accounts.&nbsp;Russian social network VK was used to record the data.</p> <p>The database is provided in SQLite3 format. The data format is the following:</p> <p><strong>&#39;sessions&#39;&nbsp;</strong>table:</p> <table> <thead> <tr> <th scope="col">Column Name</th> <th scope="col">Data Type</th> <th scope="col">Description</th> </tr> </thead> <tbody> <tr> <td>user_id</td> <td>TEXT</td> <td>Unique user&#39;s identifier.</td> </tr> <tr> <td>platform</td> <td>INTEGER</td> <td>Device identifier for the session (refer to the table below).</td> </tr> <tr> <td>time_from</td> <td>DATE</td> <td>Timestamp of the session&#39;s start.</td> </tr> <tr> <td>time_to</td> <td>DATE</td> <td>Timestamp of the session&#39;s end.</td> </tr> </tbody> </table> <p><strong>&#39;map&#39;&nbsp;</strong>table:</p> <table> <thead> <tr> <th scope="col">Column Name</th> <th scope="col">Data Type</th> <th scope="col">Description</th> </tr> </thead> <tbody> <tr> <td>user_id</td> <td>TEXT</td> <td>Unique user&#39;s identifier.</td> </tr> <tr> <td>hour</td> <td>INTEGER</td> <td>Hour from the 1st&nbsp;Jan 1970 (Unix Epoch / 3600).</td> </tr> <tr> <td>time</td> <td>INTEGER</td> <td>Accumulated online time in the hour (in seconds).</td> </tr> </tbody> </table> <p>Device identifiers:</p> <table> <tbody> <tr> <td>0</td> <td>Unknown</td> </tr> <tr> <td>1</td> <td>Web on Mobile&nbsp;</td> </tr> <tr> <td>2</td> <td>iPhone App</td> </tr> <tr> <td>3</td> <td>iPad App</td> </tr> <tr> <td>4</td> <td>Android App</td> </tr> <tr> <td>5</td> <td>Windows Phone App</td> </tr> <tr> <td>6</td> <td>Windows App</td> </tr> <tr> <td>7</td> <td>Web on Desktop</td> </tr> </tbody> </table> <p>&nbsp;</p> <p>This dataset is associated with the VKWatcher independent research project. The code used to gather the information can be found <a href="https://github.com/Azarattum/VKWatcher-Backend">on GitHub</a>.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Online Resources for Strullu-Derrien et al - The 330–320 Million-Year-Old Tranchée des Malécots (Chaudefonds-sur-Layon, South of the Armorican Massif, France): a Rare Geoheritage Site Containing In Situ Palaeobotanical Remains

<p>This repository contains the following files associate with &quot;The 330&ndash;320 Million-Year-Old Tranch&eacute;e des Mal&eacute;cots (Chaudefonds-sur-Layon, South of the Armorican Massif, France): a Rare Geoheritage Site Containing In Situ Palaeobotanical Remains&quot; by&nbsp;Christine Strullu-Derrien, Alan RT Spencer, Christopher J Cleal&nbsp;and Victor O. Leshyk.</p> <p><strong>Online Resource 1</strong> Model data as a .zip archive (301.5MB) containing .obj/.mtl and texture files for each 3D reconstruction (Models #1-4, whole site reconstruction, detailed reconstruction of the trench, and model of the mine site).</p> <p><strong>Online Resource 2</strong> Video animation showing whole site 3D model (.mp4 | 37.7MB), with quick fly-through of the Tranch&eacute;e des Mal&eacute;cots showing exposed rock and bedding of the SW wall.</p> <p><strong>Online Resource 3</strong> Video animation showing 3D Model #1 (.mp4 | 35.1MB).</p> <p><strong>Online Resource 4</strong> Video animation showing 3D Model #2 (.mp4 | 83.5MB).</p> <p><strong>Online Resource 5</strong> Video animation showing 3D Model #3 (.mp4 | 45.3MB).</p> <p><strong>Online Resource 6</strong> Video animation showing 3D Model #4 (.mp4 | 65.8MB).</p> <p><strong>Online Resource 7</strong> Video animation showing 3D model of the&nbsp;Mal&eacute;cots mine headframe (.mp4 | 14.9.0MB).</p>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Students' perceived obstacles with Forced Online Distance Learning during the CoVID-19 outbreak and their preferences to continue with the introduced teaching methods after the reopening of the University of Maribor [Project documentation]

<p>The outbreak of COVID -19 forced most universities into distance education. Three didacticians and researchers from the University of Maribor, Slovenia: Kosta Dolenc, Mateja Ploj Virtič and Andrej &Scaron;orgo formed a self-initiated initiative project group during the COVID -19 epidemic and started the first project with the working title: The Side Effects of Forced Online Distance Education (FODE).</p> <p>The aim of the second study, conducted during the first wave of the epidemic in March 2020, was to investigate the response of university students to the new situation. The project documentation provided for the Forced Online Distance Learning (FODL)&nbsp;consists of:</p> <ul> <li>abstract,</li> <li>instrument,</li> <li>copy of the descriptive statistics,</li> <li>and&nbsp;SPSS dataset.</li> </ul>

opencc-by-4.0Apr 2021View details →
zenodo44/100

Forced Continuance Intention Model of Distance Online Teaching during CoVID-19 outbreak at University of Maribor, Slovenia [Project documentation]

<p>The outbreak of COVID -19 forced most universities into distance education. Three didacticians and researchers from the University of Maribor, Slovenia: Kosta Dolenc, Mateja Ploj Virtič and Andrej &Scaron;orgo formed a self-initiated initiative project group during the COVID -19 epidemic and started the project with the working title: The Side Effects of Forced Online Distance Education (FODE).</p> <p>The aim of the first study, conducted during the first wave of the epidemic in March 2020, was to investigate the response of university teachers to the new situation. The project documentation provided for&nbsp;the Forced Online Distance Teaching (FODT) consist&nbsp;of:</p> <ul> <li>abstract,</li> <li>instrument,</li> <li>copy of the descriptive statistics, and</li> <li>SPSS dataset.</li> </ul>

opencc-by-4.0Apr 2021View details →
zenodo44/100

How does moving Public Engagement with Research Online Change Audience Diversity? Comparing Inclusion Indicators for 2019 & 2020 European Researchers' Night events

<p>Taking place annually in more than 400 cities, European Researchers&rsquo; Night is a pan- European synchronized event that aims to bring researchers closer to the public. In this paper audience profiles are compared from events in 2019 and 2020. In 2019, face-to-face events reached an estimated 1.6 million attendees, while in 2020, events shifted online due to the COVID-19 pandemic and reached an estimated 2.3 million attendees. Focusing on social inclusion metrics, survey data is analyzed across two national contexts (Ireland and Malta) in 2019 (n=656) and 2020 (n=506). The results from this exploratory, descriptive study shed light on how moving public engagement with research online shifted audience profiles. Based on prior research about the digital divide in access and use of online media, hypotheses were proposed that online European Researchers&rsquo; Night events would attract audiences with higher educational attainment levels and greater self-reported, subjective economic well-being. While changes were observed from 2019 to 2020, results for each hypothesis show a mixed picture. The first hypothesis was upheld for the highest education levels but failed for the lowest levels suggesting that the pivot to online events simultaneously attracted participants with no formal education and those with postgraduate qualifications, while attracting less of those with undergraduate or lower levels of education. The second hypothesis was not upheld, with online European Researchers&rsquo; Night events attracting audiences with slightly higher levels of economic well-being compared to face-to-face events. The findings of this study indicate that European Researchers&rsquo; Night events present a clear opportunity to measure the effects of the digital divide in relation to public engagement with research across Europe.</p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

Innovative forest products in the circular bioeconomy: online survey questionnaire and dataset

<p>This upload includes the dataset and questionnaire for an online survey with stakeholders, as part of a case study done for the BioMonitor project. The survey participants work in EU-based organizations involved in the development and manufacture of forest products, especially of the following categories: construction materials, textiles, chemicals, bioplastics, and wood-based composites.</p> <p>&nbsp;<br> <strong>About BioMonitor</strong><br> BioMonitor is an EU-funded project (biomonitor.eu) that aims to establish a sustainable and robust framework that different stakeholders can use to monitor and measure the bioeconomy and its various impacts in relation to the EU and its Member States. The BioMonitor consortium is composed of a team of universities, statistical and standardisation institutes as well as consultancies and data modelling experts.</p> <p><em>This work was supported by the BioMonitor project, which has received funding from the European Union&rsquo;s Horizon 2020 Research and Innovation Programme under Grant Agreement N&deg; 773297.</em></p>

opencc-by-4.0Jan 2022View details →
zenodo44/100

BIM online learning content in Brazil

<p>We present the result of a survey that collected three types of content on the Internet on Building Information Modeling: (i) dissemination material, (ii) training courses, and (iii) tutorials, available online for open access in Brazil or abroad.</p> <p>The identified BIM content was categorized by the following fields:</p> <ul> <li> <p>NAME: title of content;</p> </li> <li> <p>LINK: url for online location;</p> </li> <li> <p>COMPONENT OF COMPETENCE: indicates whether the component of competence is conceptual&nbsp;(theoretical knowledge) or applied &nbsp;(skill) according to Succar, Scher and Williams (2013);</p> </li> <li> <p>COMPETENCE CLASSIFICATION: level of competence (domain or executive) according to the BIMe Initiative competence table;</p> </li> <li> <p>FORMAT: website, youtube channel, or podcast;</p> </li> <li> <p>CONTENT: dissemination, tutorial, online training course;</p> </li> <li> <p>TOOL: if the content has a specific focus on a software, its name is listed;</p> </li> <li> <p>HOURS: class hours (when applicable);</p> </li> <li> <p>CERTIFICATE: yes or no (when applicable) and</p> </li> <li> <p>NOTE: explanatory text.</p> </li> </ul>

opencc-by-4.0Feb 2022View details →
zenodo44/100

Bayesian Online Learning for Energy-Aware Resource Orchestration in Virtualized RANs - Dataset

<p>Dataset providing a set of measurement of performance and power consumpetion of a virtualized Base Station (srseNB).</p>

opencc-by-4.0Jun 2021View details →
zenodo44/100

Magyar Nemzet Online

<p>This object has been created as a part of the web harvesting project of the E&ouml;tv&ouml;s Lor&aacute;nd University Department of Digital Humanities <a href="https://elte-dh.hu/en">ELTE DH</a>. Learn more about the workflow <a href="https://www.aclweb.org/anthology/2020.wac-1.5/">HERE</a> about the software used <a href="https://github.com/ELTE-DH/WebArticleCurator/">HERE</a>.The aim of the project is to make online news articles and their metadata suitable for research purposes. The archiving workflow is designed to prevent modification or manipulation of the downloaded content. The current version of the curated content with normalized formatting in standard <a href="https://tei-c.org/">TEI XML</a> format with Schema.org encoded metadata is available <a href="https://doi.org/10.5281/zenodo.6340252">HERE</a>. The detailed description of the raw content is the following:</p><ul><li>The portal&#39;s archived content (from 1998-05-24 to 2019-05-10) in WARC format available <a href="https://doi.org/10.5281/zenodo.6334896">HERE</a> (crawled: 2019-08-30T21:14:58.566822 - 2019-10-08T23:55:23.800961).</li></ul><br>Please fill in the following form before requesting access to this dataset:<a href="https://forms.office.com/r/VnP05QYTGV">ACCES FORM</a>

opencc-by-4.0Mar 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record