Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

3

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

3 results for “web search engine”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data set of the article: Using Machine Learning for Web Page Classification in Search Engine Optimization

<p>Data of investigation published&nbsp;in the article: &quot;Using Machine Learning for Web Page Classification in Search Engine Optimization&quot;</p> <p>Abstract of the article:</p> <p>This paper presents a novel approach of using machine learning algorithms based on experts&rsquo; knowledge to classify web pages into three predefined classes according to the degree of content adjustment to the search engine optimization (SEO) recommendations. In this study, classifiers were built and trained to classify an unknown sample (web page) into one of the three predefined classes and to identify important factors that affect the degree of page adjustment. The data in the training set are manually labeled by domain experts. The experimental results show that machine learning can be used for predicting the degree of adjustment of web pages to the SEO recommendations&mdash;classifier accuracy ranges from 54.59% to 69.67%, which is higher than the baseline accuracy of classification of samples in the majority class (48.83%). Practical significance of the proposed approach is in providing the core for building software agents and expert systems to automatically detect web pages, or parts of web pages, that need improvement to comply with the SEO guidelines and, therefore, potentially gain higher rankings by search engines. Also, the results of this study contribute to the field of detecting optimal values of ranking factors that search engines use to rank web pages. Experiments in this paper suggest that important factors to be taken into consideration when preparing a web page are page title, meta description, H1 tag (heading), and body text&mdash;which is aligned with the findings of previous research. Another result of this research is a new data set of manually labeled web pages that can be used in further research.&nbsp;</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

A Novel Algorithm for Estimating Web Page Ranking in Search Engine Results Pages

<p><em><strong>Abstract:</strong> </em>Search engine optimization (SEO) can make a big improvement in the traffic to a web page. Because search engines keep their main rules of ranking undeclared, it&rsquo;s important to develop models that can estimate the ranking of a web page in the search engine to be able to optimize web pages to rank higher in the search engine. The available research methodologies used machine learning algorithms to provide solutions for this target with the help of generated datasets by scraping the search engine results pages (SERP) and crawling web pages. Their proposed models suffered from the inability to be updated dynamically if the search engine updated its ranking algorithm, and their input data did not include the diversity of web pages and languages. This research will propose a novel original rank estimation algorithm that&rsquo;s able to overcome other research challenges, with a set of comparative experiments and complexity analysis. Results will show that the proposed algorithm could achieve higher values of accuracy, precision, and recall.</p> <p><strong><em>Dataset:&nbsp;</em></strong></p> <p>For research purpose, the dataset will play two roles, first, it will act the role of search engine result pages (SERP), and second, it will be used to test algorithms and calculate performance measurements.&nbsp;Dataset is consisting of 9930 web pages, aimed to identify search results pages, focusing on the top 3 pages of SERP, with 31 extracted attributes that&#39;s related to search engine optimization (SEO). The distribution of examples between class labels was balanced, with changes due to scraping operation issues, but not significantly different, with fractions of 39.9%, 34.6%, and 25.5% for the class labels page1, page2, and page 3. Feature names are: &#39;Title 1 Length&#39;, &#39;Title 2 Length&#39;, &#39;Meta Description 1 Length&#39;, &#39;Meta Description 2 Length&#39;, &#39;Meta Keywords 1 Length&#39;, &#39;H1-1 Length&#39;, &#39;H1-2 Length&#39;, &#39;H2-1 Length&#39;, &#39;H2-2 Length&#39;, &#39;Size (bytes)&#39;, &#39;Word Count&#39;, &#39;Text Ratio&#39;, &#39;Inlinks&#39;, &#39;Unique Inlinks&#39;, &#39;Unique JS Inlinks&#39;, &#39;% of Total&#39;, &#39;Outlinks&#39;, &#39;Unique Outlinks&#39;, &#39;Unique JS Outlinks&#39;, &#39;External Outlinks&#39;, &#39;Unique External Outlinks&#39;, &#39;Unique External JS Outlinks&#39;, &#39;Response Time&#39;, &#39;Status Code&#39;, &#39;Keyword in MetaDescription1&#39;, &#39;Keyword in Title1&#39;, &#39;Keyword in MetaKeywords1&#39;, &#39;Keyword in URL&#39;, &#39;Has LastModified&#39;, &#39;Keyword in Headers&#39;, and &#39;Keyword in Emphasized Text&#39;.</p> <p>The process of dataset generation involved&nbsp;scraping the search engine, extracting URLs for selected keywords, focusing on feature extraction, cleaning and preprocessing, and generating new attributes related to keywords in web pages. It&nbsp;involved also removing missing values, duplicates, and data type conversions to obtain a comprehensive dataset.<br> Keyword selection involves selecting keywords from various categories and considering diversity, including high and low traffic, long-term and short-term keywords, and generic and branded keywords. Apify online tool was used for search engine scraping with default language and US country, resulting in 388 selected keywords with 30 results per keyword. Dataset included extracted SEO features from 9991 web pages using screamingFrog desktop software and Rapidminer desktop software, determining page SEO-friendliness and comparing it to SERP rankings. Dataset cleaning involved removing redundant attributes, removing paid SERP results, replacing missing values, and converting data types. Rapidminer was used for data cleaning and preprocessing, generating new attributes related to keyword usage in web pages.<br> &nbsp;</p>

opencc-by-4.0Sep 2023View details →
zenodo32/100

User Evaluation and Metrics Analysis of a Prototype Web-based Federated Search Engine for Art and Cultural Heritage

<p>This dataset includes the quantitative data of the usage during the evaluation phase of a prototype web-based federated search engine for art and cultural heritage related content. The metrics which resulted in the dataset were in the form of a timeline of actions taken from a user (evaluator) in the course of a single session of interaction with the platform. A total of 20 different metrics were being monitored regarding the usage of the search engine, including submitting a query, a voice query, preforming a visual search, viewing a result, viewing a visual search result, updating an avatar, editing a user profile or changing user preferences, bookmarking and removing bookmarks of results and visual search results, using text to speech of all the various elements, opening the source view of a result and clicking a concept tag. All metrics included the timestamp of the event taking place and the value of the related event (e.g. the term of a search query).</p>

opencc-by-4.0Apr 2022View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record