Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

9

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

9 results for “ranking algorithm”

Learn how ShareScore rates datasets ↗
zenodo40/100

Data set of the article: Language Bias in the Google Scholar Ranking Algorithm

<p>Data of investigation published&nbsp;in the article Crist&ograve;fol Rovira; Llu&iacute;s Codina; Carlos Lopezosa&nbsp;Language Bias in the Google Scholar Ranking Algorithm. Future Internet, 2021, 13.</p> <p><strong>Abstract: </strong>The visibility of academic articles or conference papers depends on their being easily found in academic search engines, above all in Google Scholar. To enhance this visibility, search engine optimization (SEO) has been applied in recent years to academic search engines in order to optimize documents and, thereby, ensure they are better ranked in search pages (i.e., academic search engine optimization or ASEO). To achieve this degree of optimization, we first need to further our understanding of Google Scholar&rsquo;s relevance ranking algorithm, so that, based on this knowledge, we can highlight or improve those characteristics that academic documents already present and which are taken into account by the algorithm. This study seeks to advance our knowledge in this line of research by determining whether the language in which a document is published is a positioning factor in the Google Scholar relevance ranking algorithm. Here, we employ a reverse engineering research methodology based on a statistical analysis that uses Spearman&rsquo;s correlation coefficient. The results obtained point to a bias in multilingual searches conducted in Google Scholar with documents published in languages other than in English being systematically relegated to positions that make them virtually invisible. This finding has important repercussions, both for conducting searches and for optimizing positioning in Google Scholar, being especially critical for articles on subjects that are expressed in the same way in English and other languages, the case, for example, of trademarks, chemical compounds, industrial products, acronyms, drugs, diseases, etc.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Spatial Evolve Algorithm Results for Erdos Renyi Topology Median Normalized Rank - MSc Dissertation

<p>A data set containing the results of the spatial evolve lookup algorithm. The topology&nbsp;used for the spatial tournaments has been the Erdős R&eacute;nyi random network. The objective function taken into account has been the median normalized rank.&nbsp;Three files are contained here based on the list of strategies,deterministic and non, and on the sample size.&nbsp;</p>

opencc-zeroSep 2016View details →
zenodo40/100

Spatial Evolve Algorithm Results for Random Topology Median Normalized Rank - MSc Dissertation

<p>A data set containing the results of the spatial evolve lookup algorithm. The topologies used for the spatial tournaments has been random between, small world, random and complete. The objective function taken into account has been the median normalized rank. Three files are contained here based on the list of strategies,deterministic and non, and on the sample size.&nbsp;</p>

opencc-zeroSep 2016View details →
zenodo40/100

Spatial Evolve Algorithm Results for Watts Strogatz Topology Median Normalized Rank - MSc Dissertation

<p>A data set containing the results of the spatial evolve lookup algorithm. The topology&nbsp;used for the spatial tournaments has been the Watts Strogatz small world&nbsp;network. The objective function taken into account has been the median normalized rank.&nbsp;Three files are contained here based on the list of strategies,deterministic and non, and on the sample size.&nbsp;</p>

opencc-zeroSep 2016View details →
zenodo40/100

Spatial Evolve Algorithm Results for Erdős Rényi Topology Median Normalized Rank - MSc Dissertation

<p>A data set containing the results of the spatial evolve lookup algorithm. The topology&nbsp;used for the spatial tournaments has been the Erdős R&eacute;nyi random network. The objective function taken into account has been the median normalized rank. Three&nbsp;files are contained here based on the strategies list, deterministic and non and on the sample size.&nbsp;</p>

opencc-zeroSep 2016View details →
zenodo40/100

Spatial Evolve Algorithm Results for Complete Topology Median Normalized Rank - MSc Dissertation

<p>A data set containing the results of the spatial evolve lookup algorithm. The topology&nbsp;used for the spatial tournaments has been a complete&nbsp;network. The objective function taken into account has been the median normalized rank. Three files are contained here based on the list of strategies,deterministic and non, and on the sample size.&nbsp;</p>

opencc-zeroSep 2016View details →
zenodo40/100

Spatial Evolve Algorithm Results for Random Topology Minimum Normalized Rank - MSc Dissertation

<p>A data set containing the results of the spatial evolve lookup algorithm. The topologies used for the spatial tournaments has been random between, small world, random and complete. The objective function taken into account has been the minimum&nbsp;normalized rank. Two&nbsp;files are contained here based on the sample size. All 132 strategies of the Axelrod have&nbsp;been used.&nbsp;</p>

opencc-zeroSep 2016View details →
zenodo40/100

Fig. 8. Algorithmic Page Rank-Study of a Random Navigation on the Web Using Software Simulation

<p>Another diagram shows that for each site there is only one constant value independent of the<br> length of Markov Chain. Algorithmic Page Rank only depends on the number of inlinks and<br> outlinks.Table 4 contains the values for Experimental Page Rank with balanced distribution. In the<br> following figure it is shown the diagram for the values obtained for Experimental Page Rank.<br> The Experimental Page Rank values oscillate between the same limits independently of the<br> change of Markov Chain length (N).</p>

opencc-by-4.0Aug 2015View details →
zenodo36/100

A Novel Algorithm for Estimating Web Page Ranking in Search Engine Results Pages

<p><em><strong>Abstract:</strong> </em>Search engine optimization (SEO) can make a big improvement in the traffic to a web page. Because search engines keep their main rules of ranking undeclared, it&rsquo;s important to develop models that can estimate the ranking of a web page in the search engine to be able to optimize web pages to rank higher in the search engine. The available research methodologies used machine learning algorithms to provide solutions for this target with the help of generated datasets by scraping the search engine results pages (SERP) and crawling web pages. Their proposed models suffered from the inability to be updated dynamically if the search engine updated its ranking algorithm, and their input data did not include the diversity of web pages and languages. This research will propose a novel original rank estimation algorithm that&rsquo;s able to overcome other research challenges, with a set of comparative experiments and complexity analysis. Results will show that the proposed algorithm could achieve higher values of accuracy, precision, and recall.</p> <p><strong><em>Dataset:&nbsp;</em></strong></p> <p>For research purpose, the dataset will play two roles, first, it will act the role of search engine result pages (SERP), and second, it will be used to test algorithms and calculate performance measurements.&nbsp;Dataset is consisting of 9930 web pages, aimed to identify search results pages, focusing on the top 3 pages of SERP, with 31 extracted attributes that&#39;s related to search engine optimization (SEO). The distribution of examples between class labels was balanced, with changes due to scraping operation issues, but not significantly different, with fractions of 39.9%, 34.6%, and 25.5% for the class labels page1, page2, and page 3. Feature names are: &#39;Title 1 Length&#39;, &#39;Title 2 Length&#39;, &#39;Meta Description 1 Length&#39;, &#39;Meta Description 2 Length&#39;, &#39;Meta Keywords 1 Length&#39;, &#39;H1-1 Length&#39;, &#39;H1-2 Length&#39;, &#39;H2-1 Length&#39;, &#39;H2-2 Length&#39;, &#39;Size (bytes)&#39;, &#39;Word Count&#39;, &#39;Text Ratio&#39;, &#39;Inlinks&#39;, &#39;Unique Inlinks&#39;, &#39;Unique JS Inlinks&#39;, &#39;% of Total&#39;, &#39;Outlinks&#39;, &#39;Unique Outlinks&#39;, &#39;Unique JS Outlinks&#39;, &#39;External Outlinks&#39;, &#39;Unique External Outlinks&#39;, &#39;Unique External JS Outlinks&#39;, &#39;Response Time&#39;, &#39;Status Code&#39;, &#39;Keyword in MetaDescription1&#39;, &#39;Keyword in Title1&#39;, &#39;Keyword in MetaKeywords1&#39;, &#39;Keyword in URL&#39;, &#39;Has LastModified&#39;, &#39;Keyword in Headers&#39;, and &#39;Keyword in Emphasized Text&#39;.</p> <p>The process of dataset generation involved&nbsp;scraping the search engine, extracting URLs for selected keywords, focusing on feature extraction, cleaning and preprocessing, and generating new attributes related to keywords in web pages. It&nbsp;involved also removing missing values, duplicates, and data type conversions to obtain a comprehensive dataset.<br> Keyword selection involves selecting keywords from various categories and considering diversity, including high and low traffic, long-term and short-term keywords, and generic and branded keywords. Apify online tool was used for search engine scraping with default language and US country, resulting in 388 selected keywords with 30 results per keyword. Dataset included extracted SEO features from 9991 web pages using screamingFrog desktop software and Rapidminer desktop software, determining page SEO-friendliness and comparing it to SERP rankings. Dataset cleaning involved removing redundant attributes, removing paid SERP results, replacing missing values, and converting data types. Rapidminer was used for data cleaning and preprocessing, generating new attributes related to keyword usage in web pages.<br> &nbsp;</p>

opencc-by-4.0Sep 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record