Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
22
datasets available to search
ShareScore release 0.9.0
Dataset results
22 results for “web page”
Web Experience in Mobile Networks: Lessons from Two Million Page Visits
<p>Measuring and characterizing web page performance is a challenging task.</p> <p>When it comes to the mobile world, the highly varying technology characteristics coupled with the opaque network configuration make it even more difficult.</p> <p>Aiming at reproducibility, we present a large scale measurements study of web page performance collected in eleven commercial mobile networks spanning four countries.</p> <p>We build a dataset of nearly two million web browsing sessions to we shed light on the impact of different web protocols, browsers, and mobile technologies on the web performance.</p> <p>We find that the impact of mobile broadband access is sizeable.</p> <p>For example, the median page load time using mobile broadband increases by a third compared to wired access.</p> <p>Mobility clearly stresses the system, with handover causing the most evident performance penalties.</p> <p>Contrariwise, our measurements show that the adoption of HTTP/2 and QUIC has practically negligible impact.</p> <p>Our work highlights the importance of large-scale measurements.</p> <p>Even with our controlled setup, the complexity of the mobile web ecosystem is challenging to untangle.</p> <p>For this, we are releasing the dataset as open data for validation and further research.</p> <p>We also release together with the datasets we collected the scripts we use to produce the analysis we present in the paper. Please use plot_all.sh script to generate the plots in the paper, using the separate scripts from the "scripts" archive. </p> <p>Should you use any of these resources, please also make an attribution using the following reference (provided here in bibtex format):</p> <pre>@inproceedings{rajiullah2019web, title={{Web Experience in Mobile Networks: Lessons from Two Million Page Visits}}, author={Rajiullah, Mohammad and Lutu, Andra and Khatouni, Ali Safari and Fida, Mah-Rukh and Mellia, Marco and Brunstrom, Anna and Alay, Ozgu and Alfredsson, Stefan and Mancuso, Vincenzo}, booktitle={The World Wide Web Conference}, pages={1532--1543}, year={2019}, organization={ACM}, address = {San Francisco, CA, USA}, keywords = {Web Experience, HTTP2, QUIC, TCP, Mobile Broadband, Measurements} }</pre>
zdravniki.sledilnik.org web page statistics
<p>The data uploaded here were used in a conference (MI'22) submission: Kaj se skriva v ozadju zdravniki.sledilnik.org created by the same authors.</p> <p>Data is automatically created by Google and Github actions. </p> <p>Github repositories linked to the project are:</p> <ul> <li><a href="https://github.com/sledilnik/zdravniki">sledilnik/zdravniki</a></li> <li><a href="https://github.com/sledilnik/zdravniki-data">sledilnik/zdravniki-data</a></li> </ul> <p>You can read more about the project on <a href="https://zdravniki.sledilnik.org/en/">zdravniki.sledilnik.org</a>.</p> <p>Authors would like to thank all of the <a href="https://covid-19.sledilnik.org/en/stats">COVID-19 tracker</a> project members for making this possible.</p>
Automated prediction of visual complexity of web pages: Tools and evaluations
<p>This dataset includes the screenshots of the web pages used for the evaluation of ViCRAM which is described in the following paper:</p> <p>Eleni Michailidou, Sukru Eraslan, Yeliz Yesilada, and Simon Harper. 2020. Automated Prediction of Visual Complexity of Web Pages: Tools and Evaluations. International Journal of Human-Computer Studies (SCI-E, SSCI), 145, 102523.</p>
Data set of the article: Using Machine Learning for Web Page Classification in Search Engine Optimization
<p>Data of investigation published in the article: "Using Machine Learning for Web Page Classification in Search Engine Optimization"</p> <p>Abstract of the article:</p> <p>This paper presents a novel approach of using machine learning algorithms based on experts’ knowledge to classify web pages into three predefined classes according to the degree of content adjustment to the search engine optimization (SEO) recommendations. In this study, classifiers were built and trained to classify an unknown sample (web page) into one of the three predefined classes and to identify important factors that affect the degree of page adjustment. The data in the training set are manually labeled by domain experts. The experimental results show that machine learning can be used for predicting the degree of adjustment of web pages to the SEO recommendations—classifier accuracy ranges from 54.59% to 69.67%, which is higher than the baseline accuracy of classification of samples in the majority class (48.83%). Practical significance of the proposed approach is in providing the core for building software agents and expert systems to automatically detect web pages, or parts of web pages, that need improvement to comply with the SEO guidelines and, therefore, potentially gain higher rankings by search engines. Also, the results of this study contribute to the field of detecting optimal values of ranking factors that search engines use to rank web pages. Experiments in this paper suggest that important factors to be taken into consideration when preparing a web page are page title, meta description, H1 tag (heading), and body text—which is aligned with the findings of previous research. Another result of this research is a new data set of manually labeled web pages that can be used in further research. </p>
PDF exports of pages from the Smart Data Factory web site
<p>The attached zip file contains (in all three languages of the web site, namely Italian, German, and English) the following PDF exports:</p> <ol> <li>Smart Data Factory - Main page.pdf: the main page of the Smart Data Factory web site containing general information about the Smart Data Factory as well as links to offers, past projects, and the research groups of the Faculty of Computer Science of the Free University of Bozen-Bolzano</li> <li>Smart Data Factory - Offers.pdf: an example of a list of offers.</li> <li>Smart Data Factory - Offer.pdf: an example of one offer.</li> <li>Smart Data Factory - Archive items.pdf: an example of a list of archive items (past projects).</li> <li>Smart Data Factory - Archive item.pdf: an example of an archive item.</li> <li>Smart Data Factory - Research group.pdf: an example of the presentation of a research group.</li> </ol>
Figure 2-Validation of a Web Application by Using a Limited Number of Web Pages
<p>Simplifying the method of verification and testing of the web applications and maximising<br> the quality of the verification is the principal objective for the time being. We consider that the<br> method we describe in the previous sections is an important step in this matter and, by combining<br> the present techniques of verification and testing with the mode of reducing the objects that need to<br> be tested we can obtain very good results in obtaining web application of very good quality that<br> function correctly. We believe that this idea of selecting certain components from a web application<br> can be developed by using other methods of comparison among the components (that can use notion<br> as those introduces in [3], [4], [7]).</p>
Figure 2. Ranking datasets by Page Rank-Data Conflict Resolution among Same Entities in Web of Data
<p>As depicted in Figure 2 DBpedia is the top ranked data set while GeoLinked Data and<br> Eurostat are the low ranked data sets. In this stage the data sets whose rank scores are less than half<br> of the rank scores belonging to the top ranked data set are removed from the assessment.</p>
Figure 9. Experimental Page Rank dependency on Markov Chain length with balanced distribution-Study of a Random Navigation on the Web Using Software Simulation
<p>This paper explored different implementation choices for analyzing the most important<br> parameters about a web. Many researchers explored the use of new search engines for studying the<br> evolution of the web (Ntoulas, Cho and Olston, 2004). Another important research is realized about<br> the Link Structure Graph (LSG). The LSG captures a complete hyperlink structure from the web<br> and models link associations reflected in the page layout (Rodrigues, Milic-Frayling and Fortuna,<br> 2007). For further works ideas like extrapolation methods for accelerating page rank calculation can<br> be developed (Kamvar et al., 2003).</p>
Figure 7. Experimental Page Rank dependency on Markov Chain length-Study of a Random Navigation on the Web Using Software Simulation
<p>The next diagram proves that the values for Experimental Page Rank depend on the length<br> of the Markov Chain, while Algorithmic Page Rank remains constant.</p>
Fig. 8. Algorithmic Page Rank-Study of a Random Navigation on the Web Using Software Simulation
<p>Another diagram shows that for each site there is only one constant value independent of the<br> length of Markov Chain. Algorithmic Page Rank only depends on the number of inlinks and<br> outlinks.Table 4 contains the values for Experimental Page Rank with balanced distribution. In the<br> following figure it is shown the diagram for the values obtained for Experimental Page Rank.<br> The Experimental Page Rank values oscillate between the same limits independently of the<br> change of Markov Chain length (N).</p>
Replication Data for "Exploring Genetic Improvement of the Carbon Footprint of Web Pages"
<p>## Overview</p> <p>In this study, we explore automated reduction of the carbon footprint of web pages through genetic improvement, a process that produces alternative versions of a program by applying program transformations intended to optimize qualities of interest. We introduce a prototype tool that imposes transformations to HTML, CSS, and JavaScript code, as well as image resources, that minimize the quantity of data transferred and memory usage while also minimizing impact to the user experience (measured through loading time and number of changes imposed).</p> <p>In an evaluation, our tool outperforms two baselines---the original page and randomized changes---in the average case on all projects for data transfer quantity, and 80% of projects for memory usage and load time, often with large effect size. Our results illustrate the applicability of genetic improvement to reduce the carbon footprint of web components, and offer lessons that can benefit the design of future tools.</p> <p>## Data Contained in This Package</p> <p>- experiment_data/Subject Project-XX-X.xlsx</p> <p>Each spreadsheet contains data collected as part of our experiments, including the fitness scores of the final solutions.</p>
A dataset of late 1990s and early 2000s web banner ads on Chinese- and English-language web pages
<p>This dataset contains information about 22,915 unique banner ad images appearing on Chinese- and English-language web pages in the late 1990s and early 2000s. The dataset is mined from 1,384,355 archived web page snapshots downloaded from the Wayback Machine, representing 77,747 unique HTTP URLs. The URLs are collected from six printed Internet directory books published in mainland China and the United States between 1999 and 2001, as part of a larger research project on Chinese-language web archiving.</p> <p><br> For each banner ad image, the dataset provides standard image metadata such as file format and dimension. The dataset also provides the original URLs of the web pages where the banner ad image was found, timestamps of the archived web page snapshots containing the image, archived URLs of the image file, and, if available, archived URLs of web pages to which the ad image is linked. Additionally, the dataset provides text data obtained from the banner ad images using optical character recognition (OCR). We expect the dataset to be useful for researchers across a variety of disciplines and fields such as visual culture, history, media studies, and business and marketing.</p>
Figure 1-Validation of a Web Application by Using a Limited Number of Web Pages
<p>Figure 1 shows the GARWA for this example.</p>
Supplementary web page for the paper "SEAL: Integrating Program Analysis and Repository Mining"
<p>This is an archive of the supplementary material for the paper “SEAL: Integrating Program Analysis and Repository Mining” including the website and dataset. The website can also be viewed here: <a href="https://se-sic.github.io/paper-SEAL/">https://se-sic.github.io/paper-SEAL/</a></p>
A Novel Algorithm for Estimating Web Page Ranking in Search Engine Results Pages
<p><em><strong>Abstract:</strong> </em>Search engine optimization (SEO) can make a big improvement in the traffic to a web page. Because search engines keep their main rules of ranking undeclared, it’s important to develop models that can estimate the ranking of a web page in the search engine to be able to optimize web pages to rank higher in the search engine. The available research methodologies used machine learning algorithms to provide solutions for this target with the help of generated datasets by scraping the search engine results pages (SERP) and crawling web pages. Their proposed models suffered from the inability to be updated dynamically if the search engine updated its ranking algorithm, and their input data did not include the diversity of web pages and languages. This research will propose a novel original rank estimation algorithm that’s able to overcome other research challenges, with a set of comparative experiments and complexity analysis. Results will show that the proposed algorithm could achieve higher values of accuracy, precision, and recall.</p> <p><strong><em>Dataset: </em></strong></p> <p>For research purpose, the dataset will play two roles, first, it will act the role of search engine result pages (SERP), and second, it will be used to test algorithms and calculate performance measurements. Dataset is consisting of 9930 web pages, aimed to identify search results pages, focusing on the top 3 pages of SERP, with 31 extracted attributes that's related to search engine optimization (SEO). The distribution of examples between class labels was balanced, with changes due to scraping operation issues, but not significantly different, with fractions of 39.9%, 34.6%, and 25.5% for the class labels page1, page2, and page 3. Feature names are: 'Title 1 Length', 'Title 2 Length', 'Meta Description 1 Length', 'Meta Description 2 Length', 'Meta Keywords 1 Length', 'H1-1 Length', 'H1-2 Length', 'H2-1 Length', 'H2-2 Length', 'Size (bytes)', 'Word Count', 'Text Ratio', 'Inlinks', 'Unique Inlinks', 'Unique JS Inlinks', '% of Total', 'Outlinks', 'Unique Outlinks', 'Unique JS Outlinks', 'External Outlinks', 'Unique External Outlinks', 'Unique External JS Outlinks', 'Response Time', 'Status Code', 'Keyword in MetaDescription1', 'Keyword in Title1', 'Keyword in MetaKeywords1', 'Keyword in URL', 'Has LastModified', 'Keyword in Headers', and 'Keyword in Emphasized Text'.</p> <p>The process of dataset generation involved scraping the search engine, extracting URLs for selected keywords, focusing on feature extraction, cleaning and preprocessing, and generating new attributes related to keywords in web pages. It involved also removing missing values, duplicates, and data type conversions to obtain a comprehensive dataset.<br> Keyword selection involves selecting keywords from various categories and considering diversity, including high and low traffic, long-term and short-term keywords, and generic and branded keywords. Apify online tool was used for search engine scraping with default language and US country, resulting in 388 selected keywords with 30 results per keyword. Dataset included extracted SEO features from 9991 web pages using screamingFrog desktop software and Rapidminer desktop software, determining page SEO-friendliness and comparing it to SERP rankings. Dataset cleaning involved removing redundant attributes, removing paid SERP results, replacing missing values, and converting data types. Rapidminer was used for data cleaning and preprocessing, generating new attributes related to keyword usage in web pages.<br> </p>
Timed Recordings: Brozzler Loading Web Pages with Embedded Videos
<p>This upload contains timed recordings of Brozzler loading web pages with embedded video(s). </p>
Dynamic web page change content detection
<p>This dataset contains 4 parts. "SimilarWeb dataset with screenshots" is created by scraping web elements, their CSS, and corresponding screenshots in three different time intervals for around 100 web pages. Based on this data, the "SimilarWeb dataset with SSIM column" is created with the target column containing the structural similarity index measure (SSIM) of the captured screenshots. This part of the dataset is used to train machine learning regression models. To evaluate approach, "Accessible web pages dataset" and "General use web pages dataset" parts of the dataset are used.</p>
GazeMining: A Dataset of Video and Interaction Recordings on Dynamic Web Pages. Labels of Visual Change, Segmentation of Videos into Stimulus Shots, and Discovery of Visual Stimuli.
<p><strong>Recording setup</strong><br> Recordings have been taken place on 12th March 2019. Gaze data has been recorded with a Tobii 4C eye tracker with Pro license at 90 Hz. Resolution of the viewport was set to 1024x768. The display had a size of 24 inches and a resolution of 1680x1050 pixels. We polled the DOM tree every 50 milliseconds for fixed elements. We recorded the Web browsing of four participants, who followed the protocol as stored under "Dataset_visual_change/Instructions.doc".</p> <p><strong>Description of the dataset</strong><br> The dataset consists of following three subsets.</p> <p><em>1. Dataset_visual_change</em><br> The recordings of each participant p1-p4 on twelve Web sites are in the corresponding directories. For each Web site, there are nine to eleven files:</p> <ul> <li><site>.json: datacast</li> <li><site>.webm: video recording</li> <li><site>.features.csv: computer-vision features per observation</li> <li><site>.features_meta.csv: meta information about features</li> <li><site>.labels-l<X>.csv: labels of observations</li> <li><site>_meta.csv: meta information about recording</li> <li><site>_scroll_cache.csv: cache of estimated scrolling</li> <li><site>_scroll_cache_map.csv: mapping of observations to scroll cache entries</li> <li><site>_times.csv: timestamps of frames in the video recording</li> <li><site>_layer_pixels.csv: first row is the pixel count of root layer, second row is pixel count of all fixed elements</li> </ul> <p><em>2. Dataset_stimuli</em><br> Stimulus shots and visual stimuli computed with the framework. Value-based, edge-based, signal-based, and SIFT-based features have been used. The labels of the first participant's session had been used to train a random forest classifier with 100 trees for visual change classification, using the named features. The discovery has been performed on each Web site from the dataset and<br> the results are placed in the respective directories. Inside each directory, there is one directory for the detected shots and one for the discovered stimuli. In the shots directory, there is one overview as <participant>_<site>.csv file. For each shot, there are four further files:</p> <ul> <li><participant>_<site>_<shot>.png: stitched frame of the stimulus shot</li> <li><participant>_<site>_<shot>-blind.csv: frames from animations that are not contributing to the stitched frame</li> <li><participant>_<site>_<shot>-gaze.csv: gaze data (in stitched frame space)</li> <li><participant>_<site>_<shot>-mouse.csv: mouse data (in stitched frame space)</li> </ul> <p>The shots have been merged to stimuli, which are placed in the stimuli directory. The stimuli are grouped per layer (scrollable, fixed elements, etc.) and meta information is available in <layer_index>-<xpath>-meta.csv files. Furthermore, there are directories per layer, storing the discovered stimuli. Each discovered visual stimulus is represented by four files:</p> <ul> <li><stimulus_id>.png: stitched frame of the visual stimulus</li> <li><stimulus_id>-gaze.csv: gaze data (in stitched frame space)</li> <li><stimulus_id>-mouse.csv: mouse data (in stitched frame space)</li> <li><stimulus_id>-shots.csv: contained stimulus shots</li> </ul> <p><em>3. Dataset_evaluation</em><br> We have performed two evaluations of the visual stimuli discovery. One computational estimating the quality of stimuli. One case-study of an expert's task. There are two respective directories with the annotation data.</p> <p><strong>Changelog</strong><br> [1.0.2] Add counts of layer pixels per participant.<br> [1.0.1] Change to CC0 license.<br> [1.0.1] Add labels of third annotator "l3".<br> [1.0.0] Initial release.</p>
Personal Web Page In Clinical Trial Participant Education
ClinicalTrials.gov study NCT03887091. IPD Sharing: YES. Countries: 1. Publications: 0.
Companion Web Page for "Empirical Evidence of Large-Scale Diversity in API Usage of Object-Oriented Software"
<p>dataset-jars-scam2013.zip:</p> <ul> <li>Content: 3 418 Jar files, which include 382 774 different types (classes or interfaces).</li> <li>Source: All Jar files present on a machine used for performing software mining experiments for 7 years.</li> </ul> <p>type-usages-20130622.zip:</p> <ul> <li>Content: All the type-usages, filtered by type of the type-usage</li> </ul> <p>Software:</p> <ul> <li>Software to extract type-usages: <a href="https://github.com/monperrus/typeusage">https://github.com/monperrus/typeusage</a></li> </ul> <p> </p> <p>Paper:</p> <pre>@inproceedings{Mendez2013, title = {Empirical Evidence of Large-Scale Diversity in API Usage of Object-Oriented Software}, author = {Mendez, Diego and Baudry, Benoit and Monperrus, Martin}, url = {<a href="https://hal.archives-ouvertes.fr/hal-00844753/file/article.pdf">https://hal.archives-ouvertes.fr/hal-00844753/file/article.pdf</a>}, booktitle = {{International Conference on Source Code Analysis and Manipulation (SCAM'2013)}}, pages = {10}, year = {2013}, doi = {<a href="https://doi.org/10.1109/SCAM.2013.6648183">10.1109/SCAM.2013.6648183</a>}, } </pre>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.