Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
44
datasets available to search
ShareScore release 0.9.0
Dataset results
44 results for “google scholar”
Papers on Google Scholar using "sonification, auditory display, audification, sonify" as search terms
<p>Data set from a Google Scholar search in January 2023 on the terms "sonification, auditory display, audification, sonify" and added abstracts from various online ressources and keywords (automatically extracted from the abstracts only), containing:</p> <ul> <li>their title,</li> <li>a website/URL (as referenced by Google scholar),</li> <li>author(s),</li> <li>publisher information,</li> <li>their google rank in our search,</li> <li>publication year,</li> <li>the number of citations;</li> <li>paper abstracts;</li> <li>keywords generated from abstracts.</li> </ul>
Mobile Cloud Computing Bibliographic Results from Google Scholar
<p>This dataset contains all the results for the term "Mobile Cloud Computing" on Google Scholar until June 2018. The data was acquired using Publish or Perish. The data has been cleaned such that the wrong and invalid results have been removed, duplicates have been removed. Titles are accurate and fine but authors and publishers info. etc. is still unclean. For textual analysis based on paper titles, this dataset is fine. For any other factor, such as institutional or journal or authorship analysis, this isn't a good choice. </p>
Mobile Edge Computing Bibliographic Results from Google Scholar
<p>This dataset contains all the results for the term "Mobile Edge Computing" on Google Scholar until June 2018. The data was acquired using Publish or Perish. The data has been cleaned such that the wrong and invalid results have been removed, duplicates have been removed. Titles are accurate and fine but authors and publishers info. etc. is still unclean. For textual analysis based on paper titles, this dataset is fine. For any other factor, such as institutional or journal or authorship analysis, this isn't a good choice. </p>
Edge Computing Bibliographic Results from Google Scholar
<p>This dataset contains all the results for the term "Edge Computing" on Google Scholar until June 2018. The data was acquired using Publish or Perish. The data has been cleaned such that the wrong and invalid results have been removed, duplicates have been removed. Titles are accurate and fine but authors and publishers info. etc. is still unclean. For textual analysis based on paper titles, this dataset is fine. For any other factor, such as institutional or journal or authorship analysis, this isn't a good choice. </p>
Fog Computing Bibliographic Results from Google Scholar
<p>This dataset contains all the results for the term "Fog Computing" on Google Scholar until June 2018. The data was acquired using Publish or Perish. The data has been cleaned such that the wrong and invalid results have been removed, duplicates have been removed. Titles are accurate and fine but authors and publishers info. etc. is still unclean. For textual analysis based on paper titles, this dataset is fine. For any other factor, such as institutional or journal or authorship analysis, this isn't a good choice. </p>
Datos Autores USTA Google Scholar
<p>Consolidado de los datos de las publicaciones de autores con filiación USTA (Colombia) en la plataforma Google Scholar. Parte del servicio de vigilancia tecnológica del Observatorio de Cienciometría de la Universidad Santo Tomás.</p> <p>Estrategia de búsqueda:</p> <p>“universidad santo tomas” OR “santo tomás university” OR “Univ Santo Tomas” OR “Universidad Santo Tomás”</p> <p>Herramientas: Harzing’s Publish or Peris (POP), Microsoft Excel</p>
Machine Learning Articles Extracted from Google Scholar
<p>This dataset was created as part of a web scraping practice aimed at capturing academic information from Google Scholar. It contains data on <em>Machine Learning</em> research articles, including the article's title, authors, summary, direct link, citation count, and APA reference. This dataset was collected using Python and Selenium to develop skills in web scraping tools for extracting data from websites with dynamic content.</p> <p>The dataset was generated specifically as part of an academic exercise to learn and apply web scraping techniques, without a deep analysis intent for the data obtained. This dataset is intended as a resource for learning and evaluating the methods used in web data collection.</p> <p><strong>Included Fields</strong>:</p> <ul> <li><strong>title</strong>: Title of the research article.</li> <li><strong>link</strong>: Direct link to the article.</li> <li><strong>authors</strong>: Names of the article’s authors.</li> <li><strong>description</strong>: Summary or brief description of the article.</li> <li><strong>citations</strong>: Number of times the article has been cited on Google Scholar.</li> <li><strong>APA_citation</strong>: APA-formatted citation of the article.</li> </ul> <p>This dataset was created solely for educational purposes and to demonstrate the application of web scraping techniques in a controlled environment.</p>
Relação de artigos científicos sobre Humanidades Digitais obtidos pelo software Google Scholar Crawler
<p>Resultado do processamento do Google Scholar Crawler para o termo "Humanidades Digitais". Neste arquivo já foram retirados os itens duplicados e os falsos positivos (artigos que não consideramos como sendo de Humanidades Digitais.</p>
List of articles resulting from the Google Scholar search "graph based author name disambiguation" published after 1/1/2021
<p>This dataset contains the list of articles resulting from the Google Scholar search “graph based author name disambiguation” published after 1/1/2021. The list is provided for reproducibility of the survey article “Graph-based Methods for Author Name Disambiguation: A Survey” and it was obtained using the following Python script available at <a href="https://github.com/WittmannF/sort-google-scholar">https://github.com/WittmannF/sort-google-scholar</a>:</p> <blockquote> <p>$ python sortgs.py --kw “graph based author name disambiguation” --startyear 2021</p> </blockquote> <p>The command returned the CSV file that contains the first 94 publications matching the query (articles with corrupted metadata have been excluded), each with metadata about Title, Number of Citations, and Rank. The CSV contains a column that specified which articles have been eventually selected for the survey.</p>
Data set of the article: Language Bias in the Google Scholar Ranking Algorithm
<p>Data of investigation published in the article Cristòfol Rovira; Lluís Codina; Carlos Lopezosa Language Bias in the Google Scholar Ranking Algorithm. Future Internet, 2021, 13.</p> <p><strong>Abstract: </strong>The visibility of academic articles or conference papers depends on their being easily found in academic search engines, above all in Google Scholar. To enhance this visibility, search engine optimization (SEO) has been applied in recent years to academic search engines in order to optimize documents and, thereby, ensure they are better ranked in search pages (i.e., academic search engine optimization or ASEO). To achieve this degree of optimization, we first need to further our understanding of Google Scholar’s relevance ranking algorithm, so that, based on this knowledge, we can highlight or improve those characteristics that academic documents already present and which are taken into account by the algorithm. This study seeks to advance our knowledge in this line of research by determining whether the language in which a document is published is a positioning factor in the Google Scholar relevance ranking algorithm. Here, we employ a reverse engineering research methodology based on a statistical analysis that uses Spearman’s correlation coefficient. The results obtained point to a bias in multilingual searches conducted in Google Scholar with documents published in languages other than in English being systematically relegated to positions that make them virtually invisible. This finding has important repercussions, both for conducting searches and for optimizing positioning in Google Scholar, being especially critical for articles on subjects that are expressed in the same way in English and other languages, the case, for example, of trademarks, chemical compounds, industrial products, acronyms, drugs, diseases, etc.</p>
Latin American and Caribbean journals indexed in Google Scholar Metrics
<p>Dataset from a study aiming to analyze the coverage of Latin American and Caribbean journals in Google Scholar Metrics (GSM). Data from 8,205 journals from 24 countries of the region were downloaded from Latindex database. A Python script was used for automated title search and data extraction (titles, h5-index, h5-median, URLs) in GSM. For the journals not found, a manual search was carried out, with attempts by variations of the title. It was found 3,070 journals indexed in GSM, which corresponds to 37.42% of the Latindex list. The search was performed on the 2021 edition of GSM, which considers articles published between 2016 and 2020 and citations registered until July 2021. The number of all types of documents published (productivity) in the h5-index period (2016-2020) in Scopus, Journal Citation Reports, and SciELO of 1,314 journals was also identified. </p> <p>The present dataset is the result of this study, which is under peer-review in a scientific journal. </p> <p>The dataset comprises titles, h5-index; h5-median, URLs of 3,070 publications from Latin America and the Caribbean identified in Google Scholar Metrics, and the respective editorial information of the publications was extracted from Latindex</p> <p>The original language of the content was kept, mainly Spanish in the case of editorial data from Latindex. The columns descriptors are also shown in English.</p> <p>The productivity data refer to the number of all types of documents published by the journals in the period 2016-2020. Data were extracted from the InCities Journal Citation Reports, Scopus, and SciELO Citation Index (Web of Science database).</p> <p>In this version 2, only the productivity data were changed, covering a larger number of journals (1,314) and including all types of documents. Other data are the same as in the first version (https://doi.org/10.5281/zenodo.5572873).</p> <p> </p> <p> </p> <p> </p> <pre> </pre> <p> </p>
Google Scholar search record: ?start=0&q=crayfish+%22water+chemistry%22&hl=en&as_vis=0,5&as_sdt=0,5
File generated: Search date, time, timezone: 2022-04-15 17:50:45 (Europe/London) Search parameters: All these words: crayfish None of these words: This exact word or phrase: "water chemistry" Any these words: Language: en Between these years: and Number of pages exported: 1 Starting from page: 1 Citations included: TRUE Citations included: TRUE Search only in the title: FALSE Authors: Source: GS links generated: https://scholar.google.co.uk/scholar?start=0&q=crayfish+%22water+chemistry%22&hl=en&as_vis=0,5&as_sdt=0,5
Google Scholar search record: ?start=0&q=crayfish+%22water+chemistry%22&hl=en&as_vis=0,5&as_sdt=0,5
File generated: Search date, time, timezone: 2022-04-15 17:34:33 (Europe/London) Search parameters: All these words: crayfish None of these words: This exact word or phrase: "water chemistry" Any these words: Language: en Between these years: and Number of pages exported: 1 Starting from page: 1 Citations included: TRUE Citations included: TRUE Search only in the title: FALSE Authors: Source: GS links generated: https://scholar.google.co.uk/scholar?start=0&q=crayfish+%22water+chemistry%22&hl=en&as_vis=0,5&as_sdt=0,5
Google Scholar search record using GSscraper app: 2022-04-19
File generated: Search date, time, timezone: 2022-04-19 16:54:44 (Europe/London) Search parameters: All these words: crayfish None of these words: This exact word or phrase: "" Any these words: Language: en Between these years: and Number of pages exported: 1 Starting from page: 1 Citations included: TRUE Citations included: TRUE Search only in the title: FALSE Authors: Source: GS links generated: https://scholar.google.co.uk/scholar?start=0&q=crayfish&hl=en&as_vis=0&as_sdt=2007
Datos de perfiles en Google Scholar en 2017 de Universidades en Centroamérica
<p>Datos de los perfiles de investigadores en Google Scholar de seis universidades en Centroa América identificadas en esa plataforma en 2017. La tabla contiene 13 columnas y 767 filas.</p> <p><strong>Perfiles de Instituciones:</strong></p> <ul> <li>Centro Agronómico Tropical de Investigación y Enseñanza</li> <li>Instituto Tecnológico de Costa Rica</li> <li>Universidad de Costa Rica</li> <li>Universidad del Valle de Guatemala</li> <li>Universidad Nacional Costa Rica</li> <li>Universidad Tecnológica de Panamá</li> </ul> <p><strong>Diccionario de datos:</strong></p> <ul> <li>Sexo: Género del investigador, indicado como "M" para masculino o "F" para femenino.</li> <li>Institucion: Nombre de la institución a la que está afiliado el investigador.</li> <li>Nombre: Nombre completo del investigador.</li> <li>word_key: Áreas de especialización del investigador descritas mediante palabras clave.</li> <li>url_user: Enlace a la página de Google Scholar del investigador.</li> <li>Id_user: Identificador único del usuario en Google Scholar. </li> <li>citaciones: Número total de citas de las publicaciones del investigador en Google Scholar. </li> <li>cita_2011: Número de citas que tenían las publicaciones del investigador hasta el año 2011.</li> <li>hindex: Índice h actual del investigador.</li> <li> hindex_2011: Índice h del investigador hasta el año 2011. </li> <li>index10: Número de artículos del investigador que han sido citados al menos 10 veces. </li> <li>indexi10_2011: Número de artículos del investigador que habían sido citados al menos 10 veces hasta el año 2011.</li> </ul>
Datos de universidades del mundo en Webometrics 2016 y su perfil en Google Scholar
<p>Los datos contiene una hoha llamada gs_mundo con datos extraida de la página de webometrics del 2016 con su url del perfil de google scholar.</p> <p>La estructura de datos contiene:</p> <ul> <li><span>Rank</span>: Indica la posición de la universidad en el ranking basado en el número total de citas académicas registradas en Google Scholar.</li> <li><span>University</span>: Nombre oficial de la universidad. </li> <li><span>url_GS</span>: Dirección URL que lleva a la página de Google Scholar donde se pueden ver las citas académicas de la universid.</li> <li><span>Country</span>: País donde se encuentra la sede principal de la universidad. </li> <li><span>Citations</span>: Total de citas académicas en Google Scholar para las publicaciones asociadas a la universidad.</li> </ul>
Data set of the article: Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus
<p>Data of investigation published in the article "Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus".</p> <p>Abstract of the article:</p> <p>Search engine optimization (SEO) constitutes the set of methods designed to increase the visibility of, and the number of visits to, a web page by means of its ranking on the search engine results pages. Recently, SEO has also been applied to academic databases and search engines, in a trend that is in constant growth. This new approach, known as academic SEO (ASEO), has generated a field of study with considerable future growth potential due to the impact of open science. The study reported here forms part of this new field of analysis. The ranking of results is a key aspect in any information system since it determines the way in which these results are presented to the user. The aim of this study is to analyse and compare the relevance ranking algorithms employed by various academic platforms to identify the importance of citations received in their algorithms. Specifically, we analyse two search engines and two bibliographic databases: Google Scholar and Microsoft Academic, on the one hand, and Web of Science and Scopus, on the other. A reverse engineering methodology is employed based on the statistical analysis of Spearman’s correlation coefficients. The results indicate that the ranking algorithms used by Google Scholar and Microsoft are the two that are most heavily influenced by citations received. Indeed, citation counts are clearly the main SEO factor in these academic search engines. An unexpected finding is that, at certain points in time, WoS used citations received as a key ranking factor, despite the fact that WoS support documents claim this factor does not intervene.</p>
Ranking snapshot for the query "THE ROLE OF HUMAN INTELLIGENCE IN ARTIFICIAL INTELLIGENCE" performed on Google Scholar
This deposit provides a snapshot of results obtained through a Google Scholar search query on the topic of "THE ROLE OF HUMAN INTELLIGENCE IN ARTIFICIAL INTELLIGENCE". The search was performed by Alessandro Lotta, from Unipd, on May 30, 2023 at 6:34:32 PM. The captured data includes 2 pages of search results, which have been saved in the output-data.jsonld file. In addition to the citation data, the deposit includes PNG format screenshots of the search results, allowing visual reference to the captured information. The metadata for the Research Object Crate is also included in JSON format, providing essential details about the contents. The citation snapshot presented here is generated using the Unipd Ranking Citation Tool, a tool developed by Gianmaria Silvello and Alessandro Lotta (University of Padua). This tool, accessible at https://rankingcitation.dei.unipd.it
A critical review of 'just transition' publications using Google and Google Scholar
<p>This dataset offers the list of sources included in our critical review of the term 'just transition' (in English), from 1990 to 2021 in the Global North and South Africa, using Google and Google Scholar as search engines. The publications retrieved include both peer-reviewed literature and publicly available reports and documents. Results (with full citations) are categorized by actor group and type, location, and year of publication.</p>
Google Scholar search record: ?start=0&q=lobster&hl=en&as_vis=0,5&as_sdt=0,5
File generated: Search date, time, timezone: 2022-04-15 18:57:53 (Europe/London) Search parameters: All these words: lobster None of these words: This exact word or phrase: "" Any these words: Language: en Between these years: and Number of pages exported: 1 Starting from page: 1 Citations included: TRUE Citations included: TRUE Search only in the title: FALSE Authors: Source: GS links generated: https://scholar.google.co.uk/scholar?start=0&q=lobster&hl=en&as_vis=0,5&as_sdt=0,5
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.