Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
3
datasets available to search
ShareScore release 0.9.0
Dataset results
3 results for “web search engine”
Data set of the article: Using Machine Learning for Web Page Classification in Search Engine Optimization
<p>Data of investigation published in the article: "Using Machine Learning for Web Page Classification in Search Engine Optimization"</p> <p>Abstract of the article:</p> <p>This paper presents a novel approach of using machine learning algorithms based on experts’ knowledge to classify web pages into three predefined classes according to the degree of content adjustment to the search engine optimization (SEO) recommendations. In this study, classifiers were built and trained to classify an unknown sample (web page) into one of the three predefined classes and to identify important factors that affect the degree of page adjustment. The data in the training set are manually labeled by domain experts. The experimental results show that machine learning can be used for predicting the degree of adjustment of web pages to the SEO recommendations—classifier accuracy ranges from 54.59% to 69.67%, which is higher than the baseline accuracy of classification of samples in the majority class (48.83%). Practical significance of the proposed approach is in providing the core for building software agents and expert systems to automatically detect web pages, or parts of web pages, that need improvement to comply with the SEO guidelines and, therefore, potentially gain higher rankings by search engines. Also, the results of this study contribute to the field of detecting optimal values of ranking factors that search engines use to rank web pages. Experiments in this paper suggest that important factors to be taken into consideration when preparing a web page are page title, meta description, H1 tag (heading), and body text—which is aligned with the findings of previous research. Another result of this research is a new data set of manually labeled web pages that can be used in further research. </p>
A Novel Algorithm for Estimating Web Page Ranking in Search Engine Results Pages
<p><em><strong>Abstract:</strong> </em>Search engine optimization (SEO) can make a big improvement in the traffic to a web page. Because search engines keep their main rules of ranking undeclared, it’s important to develop models that can estimate the ranking of a web page in the search engine to be able to optimize web pages to rank higher in the search engine. The available research methodologies used machine learning algorithms to provide solutions for this target with the help of generated datasets by scraping the search engine results pages (SERP) and crawling web pages. Their proposed models suffered from the inability to be updated dynamically if the search engine updated its ranking algorithm, and their input data did not include the diversity of web pages and languages. This research will propose a novel original rank estimation algorithm that’s able to overcome other research challenges, with a set of comparative experiments and complexity analysis. Results will show that the proposed algorithm could achieve higher values of accuracy, precision, and recall.</p> <p><strong><em>Dataset: </em></strong></p> <p>For research purpose, the dataset will play two roles, first, it will act the role of search engine result pages (SERP), and second, it will be used to test algorithms and calculate performance measurements. Dataset is consisting of 9930 web pages, aimed to identify search results pages, focusing on the top 3 pages of SERP, with 31 extracted attributes that's related to search engine optimization (SEO). The distribution of examples between class labels was balanced, with changes due to scraping operation issues, but not significantly different, with fractions of 39.9%, 34.6%, and 25.5% for the class labels page1, page2, and page 3. Feature names are: 'Title 1 Length', 'Title 2 Length', 'Meta Description 1 Length', 'Meta Description 2 Length', 'Meta Keywords 1 Length', 'H1-1 Length', 'H1-2 Length', 'H2-1 Length', 'H2-2 Length', 'Size (bytes)', 'Word Count', 'Text Ratio', 'Inlinks', 'Unique Inlinks', 'Unique JS Inlinks', '% of Total', 'Outlinks', 'Unique Outlinks', 'Unique JS Outlinks', 'External Outlinks', 'Unique External Outlinks', 'Unique External JS Outlinks', 'Response Time', 'Status Code', 'Keyword in MetaDescription1', 'Keyword in Title1', 'Keyword in MetaKeywords1', 'Keyword in URL', 'Has LastModified', 'Keyword in Headers', and 'Keyword in Emphasized Text'.</p> <p>The process of dataset generation involved scraping the search engine, extracting URLs for selected keywords, focusing on feature extraction, cleaning and preprocessing, and generating new attributes related to keywords in web pages. It involved also removing missing values, duplicates, and data type conversions to obtain a comprehensive dataset.<br> Keyword selection involves selecting keywords from various categories and considering diversity, including high and low traffic, long-term and short-term keywords, and generic and branded keywords. Apify online tool was used for search engine scraping with default language and US country, resulting in 388 selected keywords with 30 results per keyword. Dataset included extracted SEO features from 9991 web pages using screamingFrog desktop software and Rapidminer desktop software, determining page SEO-friendliness and comparing it to SERP rankings. Dataset cleaning involved removing redundant attributes, removing paid SERP results, replacing missing values, and converting data types. Rapidminer was used for data cleaning and preprocessing, generating new attributes related to keyword usage in web pages.<br> </p>
User Evaluation and Metrics Analysis of a Prototype Web-based Federated Search Engine for Art and Cultural Heritage
<p>This dataset includes the quantitative data of the usage during the evaluation phase of a prototype web-based federated search engine for art and cultural heritage related content. The metrics which resulted in the dataset were in the form of a timeline of actions taken from a user (evaluator) in the course of a single session of interaction with the platform. A total of 20 different metrics were being monitored regarding the usage of the search engine, including submitting a query, a voice query, preforming a visual search, viewing a result, viewing a visual search result, updating an avatar, editing a user profile or changing user preferences, bookmarking and removing bookmarks of results and visual search results, using text to speech of all the various elements, opening the source view of a result and clicking a concept tag. All metrics included the timestamp of the event taking place and the value of the related event (e.g. the term of a search query).</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.