Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

2,609

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

2,609 results for “web”

Learn how ShareScore rates datasets ↗
zenodo44/100

CanVaxKB: A Web-based Cancer Vaccine Knowledgebase

<p>CanVaxKB is a web-based cancer vaccine knowledgebase. CanVaxKB collects, annotates and analyzes various types of cancer vaccines around the world. Currently it contains all cancer vaccines stored in the VIOLIN vaccine database. CanVaxKB also provides a user-friendly web interface for users to interactively search, compare, and analyze different cancer vaccines. The CanVaxKB website is here: https://violinet.org/canvaxkb.&nbsp;&nbsp;</p> <p>The Vaccine Ontology (VO) also includes the CanVaxKB stored cancer vaccine information, which is accessible at: https://github.com/vaccineontology/VO.&nbsp;</p> <p>The four supplemental files provided here are for the NCI Cancer paper about CanVaxKB. The citation for the CanVaxKB NCI Cancer is here:</p> <p>Eliyas Asfaw*, Asiyah Yu Lin*, Anthony Huffman*, Siqi Li*, Madison George*, Chloe Darancou, Madison Kalter, Nader Wehbi, Davis Bartels, Elyse Fleck, Nancy Tran, Daniel Faghihnia, Kimberly Berke, Ronak Sutariya, Farah Reyal, Youssef Tammam, Bin Zhao, Edison Ong, Zuoshuang Xiang, Virginia He, Justin Song, Andrey I. Seleznev, Jinjing Guo, Yuanyi Pan, Jie Zhang, Yongqun He. CanVaxKB: A Web-based Cancer Vaccine Knowledgebase. NCI Cancer. In press.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Replication package and appendixes for Causal inference of server- and client-side code smells in web apps evolution

<p>-Analysis&nbsp;<br>--R scripts used to make the analisys, divided by folders<br>--Data folders used in the questions</p> <p>-Appendixes - used in the article to shwo extra tables and plots</p> <p>-data folders - Aggregation of data, each app has two files, CSV and xls</p> <p>-separated data folders - 5 files for each app, with lines corresponding to the each released official version<br>--serversmells<br>--clientsmells<br>--javascriptsmells<br>--Cloc(metrics)<br>--version (all oficial releases)</p> <p>-issues_bugs<br>--data -issues by app by release&nbsp;<br>--data_bugs_more - the same but only bugs, by app by release<br>--scripts - scrips used to aggregate issues (from daily issues to by release) anf the same for bugs</p> <p>&nbsp;</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

User study data: Nudges to Mitigate Confirmation Bias during Web Search for Opinion Formation, automatic vs. reflective study

<p>Data of two user studies (282 and 307 participants), investigating the risks and benefits of warning labels with and without obfuscations to mitigate confirmation bias during web search on debated topics.</p> <p>&nbsp;</p> <p>Study Variables (study 1 and study 2)</p> <p>&nbsp;</p> <p>&nbsp;display_con: Search result display<br>&nbsp; &nbsp; - Study 1<br>&nbsp; &nbsp; &nbsp; &nbsp; - 1: targeted warning label with obfuscation<br>&nbsp; &nbsp; &nbsp; &nbsp; - 2: &nbsp;random warning label with obfuscation<br>&nbsp; &nbsp; &nbsp; &nbsp; - 3: regular (no intervention)<br>&nbsp; &nbsp; - Study 2<br>&nbsp; &nbsp; &nbsp; &nbsp; - 1: targeted warning label with obfuscation<br>&nbsp; &nbsp; &nbsp; &nbsp; - 2: targeted warning label without obfuscation<br>&nbsp; &nbsp; &nbsp; &nbsp; - 3: random warning label with obfuscation<br>&nbsp; &nbsp; &nbsp; &nbsp; - 4: random warning label without obfuscation<br>&nbsp; &nbsp; &nbsp; &nbsp; - 5: regular (no intervention)<br>- CRT_cat: Cognitive reflection<br>&nbsp; &nbsp; &nbsp; &nbsp; - 1: intuitive<br>&nbsp; &nbsp; &nbsp; &nbsp; - 2: analytic<br>- topic: Assigned debated topic<br>&nbsp; &nbsp; &nbsp; &nbsp; - 1: Is drinking milk healthy for humans?&nbsp;<br>&nbsp; &nbsp; &nbsp; &nbsp; - 2: Is homework beneficial?<br>&nbsp; &nbsp; &nbsp; &nbsp; - 3: Should people become vegetarian?<br>&nbsp; &nbsp; &nbsp; &nbsp; - 4: Should students have to wear school uniforms?<br>- clicksup_prop: Clicks on attitude-confirming (AC) search results (proportion of all clicks)<br>- clickwarn_prop: Clicks on warning label (WL) search results (proportion of all clicks)<br>- show_clicked: Clicks on show-button (number of clicks, only in conditions with obfuscation)<br>- accuracy_bias: Accuracy bias estimation (Difference between a) observed bias (as the proportion of attitude-confirming clicks) and b) perceived bias (reported in the post-interaction questionnaire and re-coded into values from 0 to 1), positive values indicate an overestimation of bias)<br>- att_change: Attitude change (Difference between attitude reported in the pre-interaction questionnaire and the post-interaction questionnaire. Negative values indicate an attitude change in the attitude-opposing direction, while positive values indicate an attitude strengthening in the attitude-supporting direction.)<br>- knowledge_1: Self-reported prior knowledge (Reported on a seven-point Likert scale ranging from non-existent to excellent as a response to how they would describe their knowledge on the topic they were assigned to)<br>- N_clicks: Cumulative clicks (Number of all clicks on search results)<br>- NFC: Need for Cognition (Mean response to 4-item subset of the NFC questionnaire)<br>- UX_usability: Usability (Mean of responses on a seven-point Likert scale to the module "usability"from the meCUE 2.0 questionnaire)<br>- UX_usefulness: Usefulness (Mean of responses on a seven-point Likert scale to the module "usefulness"from the meCUE 2.0 questionnaire)</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Base de datos de montañas y elevaciones publicada en el sitio web Summit Post (https://summitpost.org/).

<p>Lista de monta&ntilde;as y elevaciones publicadas en Summit Post (<a href="https://summitpost.org/">https://summitpost.org/</a>) hasta el 8 de noviembre de 2021.</p>

opencc-by-4.0Nov 2021View details →
zenodo44/100

Iphones publicados en la página web de Back Market el 25/03/2022 extraídos con Web Scraping

<p>El dataset recoge todos los Iphone en venta a la p&aacute;gina web de Back Market el d&iacute;a 25/03/2022 a las 19:56, correspondiendo a una extracci&oacute;n con Web Scraping de esta plataforma online, que se dedica a vender productos reacondicionados.</p> <p>El conjunto de datos incluye todos los Iphone disponibles en la tienda con sus caracter&iacute;sticas principales: nombre&nbsp;del producto (<em>Producte</em>), su capacidad en Gigas (<em>Capacitat</em>), su color (<em>Color</em>), si es libre de operador (<em>Operador</em>), el precio del producto (<em>Preu</em>), la puntuaci&oacute;n del producto&nbsp;(<em>Puntuacio</em>), la empresa que ha reacondicionado el producto (<em>Reacondicionador</em>), desde donde se env&iacute;a el producto (<em>Origen_enviament</em>), los meses de garant&iacute;a (<em>Garantia</em>) y la url de cada producto (<em>Url</em>).&nbsp;</p> <p>&nbsp;</p>

opencc-by-sa-4.0Mar 2022View details →
zenodo44/100

Code and data associated with: Searching the web builds fuller picture of arachnid trade

<p>Data and code used in the paper:&nbsp;Searching the web builds fuller picture of arachnid trade. Throughout the methods we have indicated the stage of analysis each data component was used and the code script connected. We have numbered to code and data supplements to reflect as closely as possible the order in which data generation and summary was undertaken. The following provide additional details linked to each of the data files.</p> <p>Data S1 - Website data: lang = language of the search engine used, ad hoc websites had language described after discovery; engine = the search engine used; page = the page on which the website appeared from the search engine; searchdate = search date in YYYY-mm-dd HH:MM:SS; link = link to the webpage, redacted to protect website identity; reviewdate = date revewied for arachnids being sold and search strategy; sells = whether the website sells arachnids (1 == sells); allow = whether the site explcicilt forbids automated searching (1 == allows, NA when search method was not fully automated, e.g., single page); type = the type of the website (e.g., trade, classified ads); order = whether arachnids where organised in a particular ways; target = a refined target URL to start search; method = the search method chosen, see methods for details; refine = any refinement or filter than could constrain the scope of the website to be searched; spages = the number of pages required to cycle through to cover the entire stock (also separated by ; if multiple cycles where needed or multiple single pages could be easily collected); prelimCheck = whether the website passed initial checks for arachnid selling; notes = any details that might need special attention during searches; webID = code used for subsequent data summary.</p> <p>Data S2 - Raw keyword searches outputs: species keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (only applies to Data S3); webID = the website ID.</p> <p>Data S3 &ndash; Raw keyword searches outputs: genus keywords. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID.</p> <p>Data S4 - Raw keyword search outputs: temporal sample. sp = the modern species or genus that a keyword is associated with; page = the number of the page the keyword was detected on; keyw = the exact keyword that was detected; spORgen = whether the keyword was a species binomial or just genus; termsSurrounding = the words surrounding a genus keyword detection (multiple detections separated by ;); webID = the website ID; timestamp.parse = the timestamp extracted from the archived web page; year = a simplified timestamp including only the year.</p> <p>Data S5 - LEMIS data used. An arachnid filtered version of <sup>74,75</sup>.</p> <p>Data S6 - CITES trade database data used <sup>76</sup>.</p> <p>Data S7 - CITES appendices data used <sup>77</sup>.</p> <p>Data S8 - IUCN Redlist data used <sup>78</sup>.</p> <p>Data S9 - Compiled final dataset, with data deriving from WSC, Scorpion files, ITIS, WAM and the data collection process. speciesId = a numeric code, one per species; clade = the clade the species belongs to; family = the family the species belongs to; genus = the genus of the species; species = the species epithet; author = the species authority name; year = the species authority year; parentheses = whether parentheses are needed with the authority; distribution = WSC original distribution descriptions; invalid = whether the species is considered valid; source = the species source, either World Spider Catalogue, Scorpion files, ITIS or WAM; accName = the species binomial being used as our accepted name; allNames = the accepted species binomial and all synonyms; allGenera = the accepted genus, and all other genera the species has belonged to at one point; onlineTradeSnap = whether the species was detected via a match to the accName in the snapshot data; onlineTradeSnap_Any = whether the species was detected via any synonym in the snapshot data; onlineTradeSnap_genus = whether the genus was detected via a match to the genus in the snapshot data; onlineTradeSnap_genusAny = whether the genus was detected via any synonym in the snapshot data; onlineTradeTemp = whether the species was detected via a match to the accName in the temporal data; onlineTradeTemp_Any = whether the species was detected via any synonym in the temporal data; onlineTradeTemp_genus = whether the genus was detected via a match to the genus in the temporal data; onlineTradeTemp_genusAny = whether the genus was detected via any synonym in the temporal data; onlineTradeEither = whether the species was detected via a match to the accName in the temporal data or snapshot data; onlineTradeEither_Any = whether the species was detected via any synonym in the temporal data or snapshot data; LEMIStrade = whether the species was detected via a match to the accName in the LEMIS data; LEMIStrade_Any = whether the species was detected via any synonym in the LEMIS data; LEMIStrade_genus = whether the genus was detected via any synonym in the LEMIS data; LEMIStrade_genusAny = whether the genus was detected via any synonym in the LEMIS data; CITEStrade = whether the species was detected via a match to the accName in the CITES trade database data; CITEStrade_Any = whether the species was detected via any synonym in the CITES trade database data; CITEStrade_genus = whether the genus was detected via any synonym in the CITES trade database data; CITEStrade_genusAny = whether the genus was detected via any synonym in the CITES trade database data; CITESapp = the CITES appendix the species is listed under using an exact match to the accName; CITESapp_Any = the CITES appendix the species is listed under using any match to any of the species&rsquo; synonyms; redlist = the IUCN Redlist category the species is listed under using an exact match to the accName; redlist_Any = the IUCN Redlist category the species is listed under using any match to any of the species&rsquo; synonyms; extactMatchTraded = the species is detected in any of the trade sources via a match to the accName; anyMatchTraded = the species is detected in any of the trade sources via a match to any species&rsquo; synonym.</p> <p>Data S10 - Forum listings of &ldquo;What species are you currently keeping&rdquo; from an online fora posted between 9th September 2021 and 9th October 2021, to provide an idea of online discussions. Each user with a separate list is provided in a separate tab. Morph_collector is the same as poster1, but the potential cryptic species or morphs are noted separately to make them clearer.</p> <p>Data S11 &ndash; Distribution information for spiders. Only two columns used in summaries: accName = the accepted name used throughout summaries; NAME = the country name the spider occurs in.</p> <p>Data S12 - Distribution information for scorpions. species = the accepted name used throughout summaries; NAME = the country name the scorpions occurs in.</p> <p>Code S1 - Search URL Extract.R</p> <p>Code S2 - Retrieve web data.R</p> <p>Code S3 - Temporal Classified Ads.R</p> <p>Code S4 - Keyword Generation.R</p> <p>Code S5 - Keyword Search.R</p> <p>Code S6 - LEMIS filter and summary.R</p> <p>Code S7 - Compiling results.R</p> <p>Code S8 - Summary Figures.R</p> <p>Code S9 - Temporal Figures.R</p> <p>Code S10 - New description figure.R</p> <p>Code S11 - Term exploration.R</p> <p>Code S12 - LEMIS summary and mapping.R</p>

opencc-by-4.0Dec 2021View details →
zenodo44/100

SDSS-IV Cosmic Web Catalog

<p>This repository contains the cosmic web catalog data released in the paper &quot;Cosmic Web Catalog on SDSS-IV Data with SCONCE&quot; (preparing).</p> <p>The catalog is constructed on the SDSS-IV galaxies&nbsp;and quasars (QSO) using&nbsp;our proposed Directional Subspace Constrained Mean Shift (DirSCMS) algorithm.&nbsp;We release both the cosmic filaments and local modes (i.e., local maxima of the estimated galaxy/QSO density field, which serves as candidates of galaxy clusters) within 325 thin redshift slices, each of which spans 20Mpc under the Planck15 cosmology. The entire catalog covers&nbsp;the redshift range from&nbsp;<span class="math-tex">\(z=0\)</span> to <span class="math-tex">\(z=3\)</span>.</p> <p>&nbsp; &nbsp; 1. &quot;<strong>Cosmic_filaments_2D_DirSCMS_new1</strong>&quot;: The file contains some discrete realizations of the estimated cosmic filaments in some particular redshift slices. The meaning&nbsp;of each column in the file is described&nbsp;as follows:</p> <ul> <li><strong>RA</strong> --&nbsp;right ascension.</li> <li><strong>DEC</strong> --&nbsp;declination.</li> <li><strong>z_low</strong> -- lower limit of the redshift slice.</li> <li><strong>z_high</strong>&nbsp;-- upper limit of the redshift slice.</li> <li><strong>comov_dist_low</strong> -- lower limit of the comoving distance in the redshift slice under the Planck15 cosmology.</li> <li><strong>comov_dist_high</strong>&nbsp;-- upper limit of the comoving distance in the redshift slice under the Planck15 cosmology.</li> <li><strong>bw</strong> -- smoothing bandwidth parameter for the DirSCMS algorithm in the redshift slice.</li> <li><strong>unc_meas</strong> -- uncertainty measure of the filamentary point by the nonparametric bootstrap techinque.</li> <li><strong>density</strong> -- (proportional) estimated galaxy/QSO density value at the filamentary point.</li> <li><strong>grad_Dir1</strong> -- (Riemannian) gradient of the estimated density field (first direction).</li> <li><strong>grad_Dir2</strong>&nbsp;-- (Riemannian) gradient of the estimated density field (second&nbsp;direction).</li> <li><strong>grad_Dir3</strong>&nbsp;-- (Riemannian) gradient of the estimated density field (third&nbsp;direction).</li> <li><strong>knot_label</strong> -- indicator of whether the filamentary point is a knot (i.e., the intersection of several filaments) or not.</li> </ul> <p>&nbsp; &nbsp; 2. &quot;<strong>Cosmic_local_modes_2D_DirMS_unique1</strong>&quot;:&nbsp;The file contains some discrete realizations of the estimated local modes in some particular redshift slices.&nbsp;The&nbsp;columns are&nbsp;subsumed by the ones in the cosmic filament file and has been described above.</p> <p><em>Additional notes: We provide both the &quot;csv&quot; and &quot;fits&quot; format for each of the above file.</em></p> <p>&nbsp;</p> <p>Please cite the paper when using the data in this repository.</p> <p>► Pure catalog data:</p> <p>[1]&nbsp;<strong>Cosmic Web Catalog on SDSS-IV Data with SCONCE</strong>. (In preparation)</p> <p>►Methodology:</p> <p>[1]&nbsp;Yikun Zhang, Rafael S. de Souza, and Yen-Chi Chen&nbsp;(2022). <strong>SCONCE: A Cosmic Web Finder for Spherical and Conic Geometries</strong>. <em>arXiv preprint arXiv:2207.07001</em></p> <p>[2]&nbsp;Yikun Zhang and Yen-Chi Chen (2022)&nbsp;<strong>Linear Convergence of the Subspace Constrained Mean Shift Algorithm: From Euclidean to Directional Data</strong>.&nbsp;<em>Information and Inference: A Journal of the IMA</em>, iaac005,&nbsp;<a href="https://doi.org/10.1093/imaiai/iaac005">https://doi.org/10.1093/imaiai/iaac005</a></p> <p>[3]&nbsp;Yikun Zhang and Yen-Chi Chen (2021)&nbsp;<strong>Kernel Smoothing, Mean Shift, and Their Learning Theory with Directional Data</strong>.&nbsp;<em>Journal of Machine Learning Research</em>&nbsp;<strong>22</strong>(154): 1-92.</p>

opencc-by-4.0Jun 2022View details →
zenodo44/100

Mataws annotated Web service collection

<p><strong>Description. </strong>The Mataws annotated Web service collection is a set of WS descriptions under the WSDL and OWL-S formats. It contains 816 descriptions, which were originally only syntactically described, and were annotated using our tool Mataws. Consequently, each description appears twice (once in a syntactical version, and once in a semantic version). The descriptions are also classified thematically.</p> <p>Our collection is based primarily on the FullDataset collection of the Assam project (<a href="http://www.andreas-hess.info/projects/annotator/">http://www.andreas-hess.info/projects/annotator/</a>), which we extended using WS descriptions found on the web. These individual files were classified thematically with the rest of the WSDL files, and used to assess the quality of annotation of Mataws.</p> <p><strong>Source code. </strong>The source code of our tool Mataws is available online: <a href="https://github.com/CompNet/mataws">https://github.com/CompNet/mataws</a></p> <p><strong>License.&nbsp;</strong>The annotated descriptions are shared under a Creative Commons 0 license. The original descriptions belong to their authors.</p> <p><strong>Citation. </strong>If you use our dataset, please cite the following article:</p> <ul> <li>Aksoy, C., Labatut, V., Cherifi, C. &amp; Santucci, J.-F (2011). MATAWS: A Multimodal Approach for Automatic WS Semantic Annotation. In International Conference on Networked Digital Technologies. Macau, CN : Springer. ⟨<a href="https://hal.archives-ouvertes.fr/hal-00620566">hal-00620566</a>⟩&nbsp;- DOI: <a href="https://doi.org/10.1007/978-3-642-22185-9_27">10.1007/978-3-642-22185-9_27</a></li> </ul> <p><br><code>@InProceedings{Aksoy2011,</code><br><code>&nbsp; author &nbsp; &nbsp;= {Aksoy, Cihan and Labatut, Vincent and Cherifi, Chantal and Santucci, Jean-Fran&ccedil;ois},</code><br><code>&nbsp; title &nbsp; &nbsp; = {{MATAWS}: A Multimodal Approach for Automatic WS Semantic Annotation},</code><br><code>&nbsp; booktitle = {3\textsuperscript{rd} International Conference on Networked Digital Technologies},</code><br><code>&nbsp; year &nbsp; &nbsp; &nbsp;= {2011},</code><br><code>&nbsp; volume &nbsp; &nbsp;= {136},</code><br><code>&nbsp; series &nbsp; &nbsp;= {Communications in Computer and Information Science},</code><br><code>&nbsp; pages &nbsp; &nbsp; = {319-333},</code><br><code>&nbsp; address &nbsp; = {Macau, CN},</code><br><code>&nbsp; publisher = {Springer},</code><br><code>&nbsp; doi &nbsp; &nbsp; &nbsp; = {10.1007/978-3-642-22185-9_27},</code><br><code>}</code></p>

opencc-by-4.0Jan 2015View details →
zenodo44/100

Webis-Web-Archive-Quality-22

<p>Dataset accompanying the TPDL&#39;22 publication &quot;<a href="https://webis.de/publications.html?q=Visual+Web+Archive+Quality+Assessment">Visual Web Archive Quality Assessment</a>&quot; of Theresa Elstner, Johannes Kiesel, Lars Meyer, Max Martius, Sebastian Schmidt, Benno Stein, and Martin Potthast.</p>

opencc-by-4.0Jul 2022View details →
zenodo44/100

Web requests analysis of Italy websites which use Google Analytics

<p>List of 504,038 domains of Italy found to contain Google Analytics.</p> <p>The front page for Italy-related domain names has been accessed through HTTPS or HTTP and analysed with webbkoll and jq to gather data about third-party requests, cookies and other privacy-invasive features. Together with the actual URL visited, the user/property ID is provided for 495,663 domains (extracted either from the cookies deposited or the URL of requests to Google Analytics). MX and TXT records for the domains are also provided.</p> <p>The most common ID found was 23LNSPS7Q6, with over 35k domains calling it (seemingly associated with italiaonline.it). The most common responding IP addresses were 3 AWS IPv4 addresses (over 40k domains) and 2 CloudFlare IPv6 addresses (over 12k domains).</p>

opencc-zeroJul 2022View details →
zenodo44/100

Data supporting: Combined stress of an insecticide and heatwaves or elevated temperature induce community and food web effects in a Mediterranean freshwater ecosystem

<p>Data used to obtain the results of the research paper entitled: "Combined stress of an insecticide and heatwaves or elevated temperature induce community and food web effects in a Mediterranean freshwater ecosystem", published in the journal "Water Research". The data derives from an outdoor (meso-) cosm experiment in Spain (Imdea Water, Alcala de Henares) in which the transportable temperature and heatwave control device (TENTACLE) was used to investigate the multiple stressors effects of two different climate change scenarios related to temperature (i.e., elevated temperature and reoccurring heatwaves) in combination with the neonicotinoid insecticide imidacloprid.</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Web-based Editor for Entity-Relationship Modeling with SQL Transformation Algorithm

<p>With the daily growth of data in all kinds of sectors, such as <span>Information Technolo</span><span>gies (IT)</span>, healthcare, education, commerce or telecommunication, it becomes important to use a high-performance system to manage all this data in the best possible way. Indeed, in the absence of good data management, it can be difficult for these sectors to prevent data loss, ensure safe maintenance or guarantee data security. For this reason, an effective data management system is the Entity-Relationship model. Indeed, thanks to this model, all kinds of sectors have the possibility of designing and organizing their data in relational databases, thus improving data security and better internal communication. Furthermore, it is interesting to modernize the classic approach of the Entity-Relationship model and its visual representations of data in the present day. The technologies of Augmented and Virtual Reality respond precisely to this challenge of innovation in the Entity-Relationship model. Therefore, the first objective of this thesis is to implement the Entity-Relationship model in a Meta-Modeling Platform for Augmented and Virtual Reality. The second objective is to implement an algorithm that transforms an Entity-Relationship model into SQL statements to create database tables. The Entity-Relationship model is used to design the logic of a database, and the implementation of the algorithm is used to create the database.</p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Examining LGBTQ+-related Concepts in the Semantic Web: Link Discovery, Concept Drift, Ambiguity, and Multilingual Information Reuse

<div> <h1>Examining LGBTQ+-related Concepts in the Semantic Web</h1> </div> <div> <h2>Introduction</h2> </div> <p>Welcome to the project. We study the links between LGBTQ+ ontologies and structured vocabularies. More specifically, we focus on GSSO, Homosaurus, QLIT, and Wikidata. The code is free for use with the license GPL 3,0. You can resue/extend the code for free as long as you give credits to us in your publication/data. Citation information will be added after the corresponding paper gets accepted. The paper is under submission and will be included soon.&nbsp;</p> <p>If you would like to extend this work, you may want to contact the experts in the acknowledgement before releasing your data/code about legal and ethical issues. The DOI for this version is 10.5281/zenodo.12684870. The latest code can be found at https://github.com/Multilingual-LGBTQIA-Vocabularies/Examing_LGBTQ_Concepts.&nbsp;</p> <p>To reproduce the results or extend our work, you need to take the following steps.</p> <div> <h2>Step 1: Preparing the data</h2> </div> <p>In this project, the following datasets were used:</p> <ul> <li>QLIT: version 1.0</li> <li>Homosaurus: version 3.5 and version 2.3</li> <li>Wikidata: retrieved from the SPARQL Endpoint (<a href="https://query.wikidata.org/sparql" rel="nofollow">https://query.wikidata.org/sparql</a>) and processed between 5th May and 8th May, 2024.</li> <li>GSSO: we used gsso.owl (version 2.0.10) obtained from its Github (<a href="https://github.com/Superraptor/GSSO">https://github.com/Superraptor/GSSO</a>).</li> <li>LCSH was obtained from the official website:&nbsp;<a href="https://id.loc.gov/authorities/subjects.html" rel="nofollow">https://id.loc.gov/authorities/subjects.html</a>&nbsp;on 9th May, 2024. The LCSH data was converted to its HDT format.</li> </ul> <p>Please put the corresponding files in the following folders (and change its names where necessary) to make sure that the Python scripts can find your code.</p> <ul> <li>./data/GSSO/gsso.owl</li> <li>./data/Homosaurus/v2.ttl and ./data/Homosaurus/v3.ttl</li> <li>./data/LCSH/lcsh.hdt (we used its HDT format for fast query and analysis). The original file is also attached: subjects.skosrdf.nt.</li> <li>./data/QLIT/Qlit-v1.ttl</li> </ul> <p>The case of Wikidata is more complicated. The following scripts were used for the retrival of data. These scripts are all in the folder ./data/wikidata/</p> <ul> <li>We used the Wikidata SPARQL endpoint:&nbsp;<a href="https://query.wikidata.org/" rel="nofollow">https://query.wikidata.org/</a></li> </ul> <p>The following relations from Wikidata were used while extracting triples.</p> <ul> <li>Wikidata - GSSO:&nbsp;<a href="http://www.wikidata.org/prop/direct/P9827" rel="nofollow">http://www.wikidata.org/prop/direct/P9827</a></li> <li>Wikidata - Homosaurus 2:&nbsp;<a href="http://www.wikidata.org/prop/direct/P6417" rel="nofollow">http://www.wikidata.org/prop/direct/P6417</a></li> <li>Wikidata - Homosaurus 3:&nbsp;<a href="http://www.wikidata.org/prop/direct/P10192" rel="nofollow">http://www.wikidata.org/prop/direct/P10192</a></li> <li>Wikidata - LCSH:&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a></li> </ul> <p>The generated files are:</p> <ul> <li>'wikidata-homosaurus-v2-links.nt'</li> <li>'wikidata-homosaurus-v3-links.nt'</li> <li>'wikidata-gsso-links.nt'</li> <li>'wikidata-qlit-links.nt'</li> <li>'wikidata-lcsh-links-all.nt'</li> </ul> <p>Please note that the case of Wikdiata-LCSH is more complicated: there are so many links that are nothing to do with the entities in our scope. We restrict it to only entities in the scope of this paper. See below for more details.</p> <p>You can find all the scripts in the corresponding folder in the data folder.</p> <p>All the SPARQL queries used can be found in the folder ./SPARQL/</p> <p>Note! For GSSO, the following two mistakes were corrected while preprocessing:</p> <ul> <li><a href="https://www.wikidata.org/wiki/Q1823134" rel="nofollow">https://www.wikidata.org/wiki/Q1823134</a>&nbsp;should not be used as a relation. We have replaced it with&nbsp;<a href="http://www.wikidata.org/prop/direct/P244" rel="nofollow">http://www.wikidata.org/prop/direct/P244</a>.</li> <li>Instead of referring to the page, we refer to the entity. We use&nbsp;<a href="http://www.wikidata.org/entity/" rel="nofollow">http://www.wikidata.org/entity/</a>* instead of&nbsp;<a href="https://www.wikidata.org/wiki/" rel="nofollow">https://www.wikidata.org/wiki/</a>*</li> </ul> <p>The redirection test was conducted on 30th April, 2024, between 6PM and 8PM. The files can be found in the folder of ./data/Homosaurus/redirect/.</p> <div> <h2>Integrating the data</h2> </div> <p>In the folder ./integrated_data/, you can find all the scripts related to the integrated data. Unfortunately, due to the CC-BY-NC-ND license of GSSO and Homosaurus, the integrated data will not be made available. But you can generate it with the instructions above and by using the following scripts.</p> <p>The script ./integrated_data/integrate.py takes advantage of the data generated. It first integrates a list of files of links. Then we go through the links between Wikidata and LCSH. Only those that are in the scope of the study are included.</p> <ul> <li>If your steps are correct and using the same version as we did, you should be able to get four files:</li> <li>a) the integrated file as integrated.nt</li> <li>b) the links that are relevant for this study: wikidata-lcsh-links-selected.nt.</li> <li>c) a plot of the distribution of the size of WCCs</li> <li>d) a mapping of entities and their corresponding ID of WCCs.</li> </ul> <div> <h2>Weakly Connected Components</h2> </div> <p>The weakly connected components (WCCs) were computed for the following three purposes:</p> <p>a) Discovering missing links. See the section below for details.</p> <p>b) The WCCs can be used for manual examination. These are entities that form clusters about related concepts. The intuition is that the larger they are, the more likely there is concept drift/change, ambiguity, and mistakes.</p> <p>c) Multilingual information reuse. Smaller WCCs with exactly one entity from each dataset (e.g. Homosaurus and Wikidata) can then be used to suggest labels for the one with fewer labels for some given languages. See below for more details.</p> <p>As mentioned above, the distribution has been plotted. You can find this plot here: ./integrated_data/frequency.png</p> <p>In the folder ./integrated_data/weakly_connected_components/, you can find all the WCCs and their links.</p> <p>Two examples were given in the folder. The largest WCC about sex, gender, fucking, etc. The other is about BDSM and fetish.</p> <div> <h2>Discovering missing and outdated links</h2> </div> <p>Taking advantage of WCCs, we can further find missing and outdated links. The scripts are in the folder ./discover_missing_links.</p> <p>Three examples were given. The first two is about discovering missing links. The last one is about finding outdated links.</p> <ul> <li> <p>The script ./discover_missing_links/discover_H3_LCSH.py and ./discover_missing_links/discover_QLIT_LCSH.py are scripts that outputs links that could be missing in Homosaurus and QLIT respectively. This was computed by looking at the WCCs. If two entities are both involved in the same WCC, there could be a link between them. The csv files in the same folder are the corresponding links found.</p> </li> <li> <p>The script ./discover_missing_links/find_qlit_outdated_links/ is used to discover the outdated links between QLIT and Homosaurus v3. There was only one link found.</p> </li> <li> <p>The 105 potentially missing links were taken for further review by Swedish-speaking experts from the QLIT team, which showed that 78 (72.38%) suggested links should be included: 38 (36.19%) can be included using skos:exactMatch and another 38 (36.19%) using skos:closeMatch. 28 (26.67%) suggested links are incorrect. The manual annotation are included in the file ./discover_missing_links/Annotated_found_new_links_qlit-lcsh.xlsx.</p> </li> </ul> <div> <h2>Multilingual Information Reuse</h2> </div> <p>You can find two attempts in the folders about the use of GSSO and Wikidata for Homosaurus respectively.</p> <ul> <li>./WCC-based-gsso-multilingual_info_reuse/</li> <li>./WCC-based-wikidata-multilingual_info_reuse/</li> </ul> <p>Additionally, we provide also some code for the reuse of Wikidata multilingual info for QLIT. It's in the folder</p> <ul> <li>./WCC-based-QLIT-info-reuse-from-Wikidata/</li> </ul> <p>They follow very similar steps:</p> <ol> <li> <p>Compute the one-to-one mapping using the WCCs. The script is named compute-one-to-one-mapping.py</p> </li> <li> <p>Extract the multilingual labels from sources. The corresponding file is extract_multilingual_labels_from_one_to_one_mappings.py</p> </li> <li> <p>Provide the extracted multilingual as suggestions for targeting entities. The name of the corresponding files are like "*suggesting-labels.py", where the * is replaced by the actual source/target.</p> </li> </ol> <p>For GSSO, we use the following relations:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasRelatedSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasSynonym</a></li> <li><a href="http://www.geneontology.org/formats/oboInOwl#hasExactSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasExactSynonym</a></li> <li><a href="http://purl.org/dc/terms/replaces" rel="nofollow">http://purl.org/dc/terms/replaces</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P5191" rel="nofollow">https://www.wikidata.org/wiki/Property:P5191</a></li> <li><a href="https://www.wikidata.org/wiki/Property:P1813" rel="nofollow">https://www.wikidata.org/wiki/Property:P1813</a></li> <li><a href="https://schema.org/alternateName" rel="nofollow">https://schema.org/alternateName</a></li> <li><a href="http://www.w3.org/2002/07/owl#annotatedTarget" rel="nofollow">http://www.w3.org/2002/07/owl#annotatedTarget</a></li> </ul> <p>Additioinally, we found the relation to be studied in the future:&nbsp;<a href="http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym" rel="nofollow">http://www.geneontology.org/formats/oboInOwl#hasNarrowSynonym</a></p> <p>For Wikidata, there are only two:</p> <ul> <li><a href="http://www.w3.org/2000/01/rdf-schema#label" rel="nofollow">http://www.w3.org/2000/01/rdf-schema#label</a></li> <li><a href="http://www.w3.org/2004/02/skos/core#altLabel" rel="nofollow">http://www.w3.org/2004/02/skos/core#altLabel</a></li> </ul> <div> <h2>Additional analysis</h2> </div> <p>Additionally, we perform an analysis using only redirection and replacement for GSSO and Homosaurus. The scripts are in the folder ./additional_test_gsso_multilingual_info_reuse. We consider also Homosaurus v2. This additional analysis shows the following:</p> <ul> <li> <p>For the Turkish language, in total there are 103 triples about labels about 23 entities. The average suggested labels per entity is 3.0.</p> </li> <li> <p>For the Spanish language, in total there are 205 triples about labels about 43 entities. The average suggested labels per entity is 2.12.</p> </li> <li> <p>For the French language, in total there are 277 triples about labels about 47 entities. The average suggested labels per entity is 2.19.</p> </li> <li> <p>For the Danish language, in total there are 115 triples about labels about 47 entities. The average suggested labels per entity is 2.70.</p> </li> </ul> <p>Some analysis about the replacement relations of Homosaurus is in the folder ./data/Homosaurus/replace_relations_homosaurus/.</p> <p>Finally, some additional analysis is included in the folder ./analysis_integrated_graph. Currently, there is only one that is about outdated entities in Homosaurus v3. Some more analysis will be added in the future.</p> <div> <h2>Acknowledgement</h2> </div> <p>The authors appreciate the help of the following researchers:</p> <ul> <li>Siska Humlesj&ouml;, QLIT, G&ouml;teborgs Universitet (<a href="mailto:siska.humlesjo@lir.gu.se">siska.humlesjo@lir.gu.se</a>)</li> <li>Olov Kristr&ouml;m, former member of QLIT</li> <li>Jack van der Wel, IHLIA (<a href="mailto:jack@ihlia.nl">jack@ihlia.nl</a>)</li> <li>Clair Kronk, GSSO (<a href="mailto:clair.kronk@mountsinai.org">clair.kronk@mountsinai.org</a>)</li> </ul> <div> <p>If you would like to extend this work, you may want to contact them before releasing your data/code about legal and ethical issues.</p> <h2>Contact</h2> </div> <ul> <li>Shuai Wang, Vrije Universiteit Amsterdam (<a href="mailto:shuai.wang@vu.nl">shuai.wang@vu.nl</a>)</li> <li>Maria Adamidou, Vrije Universiteit Amsterdam (<a href="mailto:m.adamidou@student.vu.nl">m.adamidou@student.vu.nl</a>)</li> </ul> <p>&nbsp;</p> <p>Thank you very much for your interest in our project!</p>

opengpl-3.0-or-laterJul 2024View details →
zenodo44/100

Supporting dataset for: Repository optimisation & techniques to improve discoverability and web impact : an evaluation

<p>This dataset supports the working paper, &quot;Repository optimisation &amp; techniques to improve discoverability and web impact : an evaluation&quot;, currently under review for publication and available as a preprint at:&nbsp;<a href="https://doi.org/10.17868/65389/">https://doi.org/10.17868/65389/</a>.&nbsp;</p> <ul> <li>Macgregor, G. (2018). <em>Repository optimisation techniques to improve discoverability and&nbsp;web impact: an evaluation</em>. (pp. 1-13). Glasgow: University of Strathclyde [Strathprints repository].&nbsp;Available: <a href="https://doi.org/10.17868/65389/">https://doi.org/10.17868/65389/</a></li> </ul> <p>The dataset comprises a single OpenDocument Spreadsheet (.ods) format file containing seven&nbsp;data sheets of data pertaining to COUNTER compliant usage statistics, search query traffic from Google Search Console, web traffic data for Google Analytics and Google Scholar, and usage statistics from IRStats2. All data relate to the EPrints repository, Strathprints, based at the University of Strathclyde.</p>

opencc-by-4.0Aug 2018View details →
zenodo44/100

A decade of Semantic Web research through the lenses of a mixed methods approach (Resources)

<p>This work has been submitted to&nbsp;<a href="http://www.semantic-web-journal.net/content/decade-semantic-web-research-through-lenses-mixed-methods-approach">Semantic Web Journal</a>. We provide here resources to reproduce our approach.</p> <p>In this paper, we aim to provide a broader and more complete picture of Semantic Web topics and trends by adopting a mixed methods methodology, which allows a combined use of both qualitative and quantitative approaches. Concretely, we build on a qualitative analysis of the main seminal papers, which adopt a top-down approach, and on quantitative results derived with three bottom-up data-driven approaches (<a href="https://technologies.kmi.open.ac.uk/Rexplore/">Rexplore</a>, <a href="http://saffron.insight-centre.org/">Saffron</a>, <a href="https://www.poolparty.biz/">PoolParty</a>), on a corpus of Semantic Web papers published in the last decade. In this process, we both use the latter for &ldquo;fact-checking&rdquo; on the former and also to derive key findings in relation to the strengths and weaknesses of top-down and bottom-up approaches to research topic identification.</p> <p>Please access the full set of resources at:&nbsp;<a href="https://aic.ai.wu.ac.at/qadlod/SW/">https://aic.ai.wu.ac.at/qadlod/SW/</a></p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Measuring Web Latency and Rendering Performance: Method, Tools & Longitudinal Dataset

<p>The dataset used in the paper entitled &quot;Measuring Web Latency and Rendering Performance: Method, Tools &amp; Longitudinal Dataset&quot; published in IEEE Transactions for Network and Service Management.&nbsp;</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

Wikimedia Commons photos by prominent users and their usage across the web

<p>Extract from the Wikimedia Commons database containing a list of users selected by the community for having uploaded high quality photos; list of 310k photos of theirs and of the subset of 59k photos sent to Infringement.Report for matching; list of domains whose matches were ignored as not useful for copyleft license enforcement. Domains were then matched for their rank in the Tranco list and the number of image usages found, and ranked by a mix of the two criteria.</p>

opencc-zeroDec 2018View details →
zenodo44/100

Ranking web of Repositories

<p>TRANSPARENT RANKING: All Repositories by Google Scholar.</p> <p>Incluye Top 10 repositorios de Colombia 2018 y 2019</p> <p>Fuente: http://repositories.webometrics.info/en/node/30</p>

opencc-by-4.0Aug 2019View details →
zenodo44/100

Web robot detection - Server logs

<p>This dataset contains server logs from the search engine of the library and information center of the Aristotle University of Thessaloniki in Greece (<a href="http://search.lib.auth.gr/">http://search.lib.auth.gr/</a>). The search engine enables users to check the availability of books and other written works, and search for digitized material and scientific publications. The server logs obtained span an entire month, from March 1st to March 31 2018 and consist of 4,091,155 requests with an average of 131,973 requests per day and a standard deviation of 36,996.7 requests. In total, there are requests from 27,061 unique IP addresses and 3,441 unique user-agent strings. The server logs are in JSON format and they are anonymized by masking the last 6 digits of the IP address and by hashing the last part of the URLs requested (after last /). The dataset also contains the processed form of the server logs as a labelled dataset of log entries grouped into sessions along with their extracted features (simple semantic features). We make this dataset publicly available, the first one in this domain, in order to provide a common ground for testing web robot detection methods, as well as other methods that analyze server logs.<br> <br> &nbsp;</p>

opencc-by-4.0Oct 2019View details →
zenodo44/100

Usability Testing Data for Web Application Prototype: Enhancing Efficiency and Transparency in Ghana's Rental Housing Market

<p><span>The dataset includes both quantitative and qualitative responses from participants who tested the web application prototype designed to enhance decision-making in Ghana's rental housing market. The testing focused on evaluating the user interface, ease of use, satisfaction levels, and the effectiveness of key functionalities.</span></p>

opencc-by-4.0Sep 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record