Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

56

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

56 results for “Scholarly Data”

Learn how ShareScore rates datasets ↗
zenodo44/100

Understanding the Publish-Review-Curate (PRC) Model of Scholarly Communication - Data and Code

<p>Summary data for the number of articles submitted to publish-review-curate platforms as of August 2024 (Figure 1) [Update 14 Nov 2024: Added JMIRx. Data still from August 2024]</p> <p>Summary data for the number of articles reviewed by review platforms (Figure 2)</p> <p>Analysis code to produce Figures 1 and 2</p> <p>Code to extract articles for inclusion in data</p>

opencc-by-4.0Oct 2024View details →
zenodo44/100

Raw and aggregated data for the study introduced in the article "An analysis of citing and referencing habits across all scholarly disciplines: approaches and trends in bibliographic metadata errors"

<p>This dataset contains all the raw data and aggregated data subject of the study introduced in the article &quot;An analysis of citing and referencing habits across all scholarly disciplines: approaches and trends in bibliographic metadata errors&quot;. The study is based on the bibliographic and citation data contained in 729 articles published in 147 journals in 27 subject areas. The articles contained a total amount of 34,140 bibliographic references and 55,100 mentions and quotations overall.</p> <p>The dataset is composed of a series of files:</p> <ul> <li>the files &quot;subject_area_&lt;discipline-name&gt;.csv&quot; contain the raw data of the articles published in the journals of all the disciplines considered in the study;</li> <li>the file &quot;article_data_summary.csv&quot; contains the aggregated data created considering the raw data in the previous files, which have been used to creating all the tables and figures in the article;</li> <li>the file &quot;starred_metadata_set.csv&quot; contains information about the most used subset of bibliographic metadata;</li> <li>the file &quot;journals_selection.csv&quot; contains information about all the journals selected for the study.</li> </ul>

opencc-zeroAug 2021View details →
zenodo44/100

Founding, Running, and Improving Scholarly Journals_Questionnaire Raw Data

<p>This is the raw data set of a questionnaire on founding, running, and improving scholarly journals and the availability of educational material for these purposes. The questionnaire was prepared on the SoSci Survey platform and the raw data downloaded from it. The questionnaire had 93 respondents who were nearly all Editors or Scholarly Publishing Professionals working with academic journals. The questionnaire was prepared for the purposes of presenting a poster at the Society for Scholarly Publishing&#39;s (SSP) Annual Conference in Portland, Oregon, USA, in May 2023.</p> <p>The Findings of the project are available here: https://zenodo.org/record/7924220#.ZFy_SaXP3cs</p> <p>A Resource Compendium is available here: https://zenodo.org/record/7924268#.ZFy_DaXP3cs</p> <p>&nbsp;</p>

opencc-by-4.0May 2023View details →
zenodo40/100

Data set of the article: Language Bias in the Google Scholar Ranking Algorithm

<p>Data of investigation published&nbsp;in the article Crist&ograve;fol Rovira; Llu&iacute;s Codina; Carlos Lopezosa&nbsp;Language Bias in the Google Scholar Ranking Algorithm. Future Internet, 2021, 13.</p> <p><strong>Abstract: </strong>The visibility of academic articles or conference papers depends on their being easily found in academic search engines, above all in Google Scholar. To enhance this visibility, search engine optimization (SEO) has been applied in recent years to academic search engines in order to optimize documents and, thereby, ensure they are better ranked in search pages (i.e., academic search engine optimization or ASEO). To achieve this degree of optimization, we first need to further our understanding of Google Scholar&rsquo;s relevance ranking algorithm, so that, based on this knowledge, we can highlight or improve those characteristics that academic documents already present and which are taken into account by the algorithm. This study seeks to advance our knowledge in this line of research by determining whether the language in which a document is published is a positioning factor in the Google Scholar relevance ranking algorithm. Here, we employ a reverse engineering research methodology based on a statistical analysis that uses Spearman&rsquo;s correlation coefficient. The results obtained point to a bias in multilingual searches conducted in Google Scholar with documents published in languages other than in English being systematically relegated to positions that make them virtually invisible. This finding has important repercussions, both for conducting searches and for optimizing positioning in Google Scholar, being especially critical for articles on subjects that are expressed in the same way in English and other languages, the case, for example, of trademarks, chemical compounds, industrial products, acronyms, drugs, diseases, etc.</p>

opencc-by-4.0Jan 2021View details →
zenodo40/100

Innovations in scholarly communication - data of the global 2015-2016 survey

<p>Innovations in scholarly communication - data of the global 2015-2016 survey.</p> <p>This data set contains:</p> <ul> <li>Full raw (anonymized) and cleaned data files of the 2015-2016 global Survey on Innovations in Scholarly Communication. Data are in xls&nbsp;format (raw and cleaned) and csv format (only for cleaned data as the raw data contain non-Roman script).</li> <li>Survey questionnaires for 7 languages (zipped PDFs)</li> <li>Variable list (xls)</li> <li>Readme file (txt)</li> </ul> <p>The data files contain &gt;3,000,000 cells, and thus cannot be opened in their entirety in Google Drive.</p> <p>Many new websites and online tools have come into existence to support<br /> scholarly communication in all phases of the research workflow. To what extent<br /> researchers are using these and more traditional tools has been largely<br /> unknown. This 2015-2016 survey aimed to fill that gap. Its results may help<br /> decision making by stakeholders supporting researchers and may also help<br /> researchers wishing to reflect on their own online workflows. In addition,<br /> information on tools usage can inform studies of changing research workflows.<br /> The online survey employed an open, non-probability sample. A largely<br /> self-selected group of 20663 researchers, librarians, editors, publishers and<br /> other groups involved in research took the survey, which was available in seven<br /> languages. The survey was open from May 10, 2015 to February 10, 2016. It<br /> captured information on tool usage for 17 research activities, stance towards<br /> open access and open science, and expectations of the most important<br /> development in scholarly communication. Respondents&rsquo; demographics<br /> included research roles, country of affiliation, research discipline and year of<br /> first publication.</p> <p>A full description of data collection, survey response and methodology is in a data publication in F1000 Research:</p> <p>Kramer, Bianca &amp; Jeroen Bosman (2016) Innovations in scholarly communication - global survey on research tool usage. F1000 Research. DOI:10.12688/f1000research.8414.1</p> <p>Contact:</p> <p>Jeroen Bosman:&nbsp;http://orcid.org/0000-0001-5796-2727 / j.bosman@uu.nl</p> <p>Bianca Kramer:&nbsp;http://orcid.org/0000-0002-5965-6560 / b.m.r.kramer@uu.nl</p>

opencc-zeroApr 2016View details →
zenodo40/100

Scholarly Wikidata: Population and Exploration of Conference Data in Wikidata using LLMs

<p>This dataset provides the input data and intermediate results of the paper titled "Scholarly Wikidata: Population and Exploration of Conference Data in Wikidata using Large Language Models and Semantic Web Techniques". It contains the following resources.</p> <ul> <li>conference proceedings front matter links - these links can be used to download the pdf files of the conference proceeding front matters that include information about the number of submitted and accepted papers that can be used to calculate acceptance rates, names of all conference organization committee members, list of programme committee and senior programme member names for each track with other interesting facts such as the main topics of the submitted papers and emerging topics according to the editors, etc.</li> <li>web crawl of conference websites - this contains a set of crawled content from each conference website in both HTML and text formats. Each file contains web pages from a specific conference along with the page URL, page title, and page content. Information such as important dates (deadlines) and other announcements can be extracted from the content of the web sites.&nbsp;</li> <li>papers and paper-authors list for each conference in a given conference series - this contains the paper list along with their corresponding authors for each conference series extracted from DBLP.&nbsp;</li> <li>OpenRefine projects - this contains examples of open refile projects that were used to perform entity linking and reconciliation as well as the schemas that was used to map the tabular data columns to Wikidata properties, and qualifiers and cell values to Wikidata entities.</li> <li>evaluation benchmark - this contains the outputs of LLM generations for the tasks (a) extracting the number of submitted and accepted papers per each track at a given conference, (b) extraction of organizers with their roles for each conference, (c) extraction of programme committee members with track and their role (member, SPC member), and (d) extraction of important dates or deadlines for each activity (submission, notification, etc.) in each track.&nbsp;</li> </ul> <p>The corresponding source code is available at the <a href="https://github.com/scholarly-wikidata/scholarly-wikidata/">scholary-data repo</a>.</p>

opencc-zeroApr 2024View details →
zenodo40/100

Analysis of scholarly repositories' availability. Data and notebooks.

<p>These datasets and companion Jupyter notebooks supplement the publication &quot;Knock knock! Who&#39;s there?&#39;&#39;&nbsp;A study on scholarly repositories&#39; availability&quot; accepted at TPDL 2022, Padova, Italy.</p>

opencc-by-4.0Jul 2022View details →
zenodo40/100

Data for "Researchers and their data. A study based on the use of the word data in scholarly articles"

<p><em>Data</em> is one of the most used terms in scientific vocabulary. This article focusses on the relationship between data and research by analyzing the contexts of occurrence of the word <em>data</em> in a corpus of 72,471 research articles (1980-2012) from two distinct fields (Social sciences, Physical sciences). The aim is to shed light on the issues raised by research on data, namely the difficulty of defining what is considered as data, the transformations that data undergo during the research process and how they gain value for researchers who hold them. Relying on the distribution of occurrences throughout the texts and over time, it demonstrates that the word <em>data </em>mostly occurs at the beginning and at the end of research articles. Adjectives and verbs accompanying the noun <em>data</em> turn out to be even more important than <em>data</em> itself in specifying data. The increase in the use of possessive pronouns at the end of the articles reveals that authors tend to claim ownership of their data at the very end of the research process. Our research demonstrates that even if data handling operations are increasingly frequent, they are still described with imprecise verbs that do not reflect the complexity of these transformations.</p>

opencc-by-4.0Jan 2022View details →
zenodo40/100

Journal data for European scholarly journals

<p>This relates to the following study: https://doi.org/10.5281/zenodo.5909512</p> <p>The methodology is described in the linked manuscript.</p> <p>&nbsp;</p>

opencc-by-4.0Oct 2022View details →
zenodo40/100

Replication data for: Bilateral flows and rates of international migration of scholars for 210 countries and areas for the period 1998-2020

<h3>Data and code for performing analyses and plotting figures for "Bilateral flows and rates of international migration of scholars for 210 countries and areas for the period 1998-2020"</h3> <p>The code and data can also be found at https://github.com/MPIDR/Global-flows-and-rates-of-international-migration-of-scholars/</p> <p><strong>Abstract</strong>: A lack of comprehensive migration data is a major barrier for understanding the causes and consequences of migration processes, including for specific groups like high-skilled migrants. We leverage large-scale bibliometric data from Scopus and OpenAlex to trace the global movements of scholars. Based on our empirical validations, we develop pre-processing steps and offer best practices for the measurement and identification of migration events. We have prepared a publicly accessible dataset that shows a high level of correlation between the counts of scholars in Scopus and OpenAlex for most countries. Although OpenAlex has more extensive coverage of non-Western countries, the highest correlations with Scopus are observed in Western countries. We share aggregated yearly estimates of international migration rates and of bilateral flows for 210 countries and areas worldwide for the period 1998-2020 and describe the data structure and usage notes. We expect that the publicly shared dataset will enable researchers to further study the causes and the consequences of migration of scholars to forecast the future mobility of global academic talent.</p>

opencc-by-4.0May 2024View details →
zenodo40/100

Data of Digital Scholarly Edition of the Diary of the Travel of Heinrich XI. Reuß-Greiz 1740-1742

Open the record for dataset details and reuse information.

opencc-by-4.0Jul 2024View details →
zenodo40/100

Network data for the paper: Intellectual and social similarity among scholarly journals.

<p>Network data used for the analysis contained in&nbsp;Baccini A, Barabesi L, Gingras Y, Kalfaoui M (2019) Intellectual and social similarity among scholarly journals. An exploratory comparison of the networks of editors, authors and co-citations.</p> <p>Data are in .net format for Pajek software</p> <p>CC indicates co-citation network.</p> <p>IA indicated Interlocking authorship network.</p> <p>IE indicates interlocking editorship network.</p> <p>Stat is for statistics; Econ is for economics; ILS is for information and library science.</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Jul 2019View details →
zenodo40/100

Data set of the article: Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus

<p>Data of investigation published&nbsp;in the article &quot;Ranking by relevance and citation counts, a comparative study: Google Scholar, Microsoft Academic, WoS and Scopus&quot;.</p> <p>Abstract of the article:</p> <p>Search engine optimization (SEO) constitutes the set of methods designed to increase the visibility of, and the number of visits to, a web page by means of its ranking on the search engine results pages. Recently, SEO has also been applied to academic databases and search engines, in a trend that is in constant growth. This new approach, known as academic SEO (ASEO), has generated a field of study with considerable future growth potential due to the impact of open science. The study reported here forms part of this new field of analysis. The ranking of results is a key aspect in any information system since it determines the way in which these results are presented to the user. The aim of this study is to analyse and compare the relevance ranking algorithms employed by various academic platforms to identify the importance of citations received in their algorithms. Specifically, we analyse two search engines and two bibliographic databases: Google Scholar and Microsoft Academic, on the one hand, and Web of Science and Scopus, on the other. A reverse engineering methodology is employed based on the statistical analysis of Spearman&rsquo;s correlation coefficients. The results indicate that the ranking algorithms used by Google Scholar and Microsoft are the two that are most heavily influenced by citations received. Indeed, citation counts are clearly the main SEO factor in these academic search engines. An unexpected finding is that, at certain points in time, WoS used citations received as a key ranking factor, despite the fact that WoS support documents claim this factor does not intervene.</p>

opencc-by-4.0Aug 2019View details →
zenodo40/100

Research Data and Software (re)Use Indications in High Energy Physics related Scholarly Works (Full dataset)

<p>This dataset contains research data and software (re)use indications (formal citations, informal mentions) in scholarly works related to High Energy Physics. 1,411 research and software indications were identified by a mix of approaches: use of citation discovery services and multiple search approaches in Google Scholar. The dataset contains indications&nbsp; by what approach the (re)use indications were found. All identified research data and software (re)use indications were classified according to their purpose, location, and elements.</p> <p>The data was collected in 2018 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

Identification of Research Data and Software (re)Use Indications in High Energy Physics related Scholarly Works (1951 - 2018)

<p>This dataset contains research data and software (re)use indications in High Energy Physics related scholarly works. A minimal random sample of scholarly works that contain highly processed research data was taken from HEPData and manually read for research data and software (re)use indications. The sample contains 368 works, which were published between 1951 and 2018.</p> <p>The data was collected in 2019 for a PhD thesis on research data and software (re)use indications in scholarly works.</p>

opencc-by-4.0Jun 2021View details →
zenodo40/100

A Novel Curated Scholarly Graph Connecting Textual and Data Publications

<p>This dataset contains an open and curated scholarly graph we built&nbsp;as a training and test set for data discovery, data connection, author disambiguation, and link prediction tasks.&nbsp;This graph represents the European Marine Science community included in the OpenAIRE Graph.&nbsp;The nodes of the graph we release&nbsp;represent publications, datasets, software, and authors respectively; edges interconnecting research products always have the publication as source, and the dataset/software as target. In addition, edges are labeled with semantics that outline whether the publication is <em>referencing, citing, documenting</em>, or <em>supplementing</em> the related outcome. To curate and enrich nodes metadata and edges semantics, we relied on the information extracted from the PDF of the publications and the datasets/software webpages respectively. We curated the authors so to remove duplicated nodes representing the same person.&nbsp;</p> <p>The resource we release counts 4,047 publications, 5,488 datasets, 22 software, 21,561 authors, and 9,692 edges connect publications to datasets/software. This graph is in the <em>curated_MES</em>&nbsp;folder. We provide this resource as:</p> <ol> <li>a property graph: we provide the dump that can be imported in neo4j</li> <li>5 jsonl files containing publications, datasets, software, authors, and relationships respectively. Each line of a jsonl file contains a JSON object representing a node and contains the&nbsp;metadata of that&nbsp;node (or a relationship).</li> </ol> <p>We provide two additional scholarly graphs:</p> <ul> <li>The curated MES graph with the removed edges. During the curation we removed some edges since&nbsp;they were labeled with an inconsistent or imprecise semantics. This graph includes the same nodes and edges as the previous one, and, in addition, it contains the edges removed during the curation pipeline; these edges are marked as <em>Removed</em>.&nbsp;This graph is in the <em>curated_MES_with_removed_semantics</em> folder.<br> &nbsp;</li> <li>The original MES community of OpenAIRE. It represents the MES community extracted from the OpenAIRE Research Graph. This graph has not been curated, and the metadata and semantics are those of the OpenAIRE Research Graph. This graph is in the <em>original_MES_community</em> folder.</li> </ul> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Dec 2022View details →
dryad36/100

Data from: TRANSCENDS: a career development program for underrepresented in medicine scholars in academic neurology

<p class="MsoNoSpacing"><b>Background: </b>The Training in Research for Academic Neurologists to Sustain Careers and Enhance the Numbers of Diverse Scholars (TRANSCENDS) program is a career advancement opportunity for individuals underrepresented in biomedical research, funded by the National Institute and Neurological Disorders and Stroke; and American Academy of Neurology (AAN).</p> <p class="MsoNoSpacing"><b>Objective:</b> To report on qualitative and quantitative outcomes in TRANSCENDS.</p> <p class="MsoNoSpacing"><b>Design:</b> Early career individuals (neurology fellows and junior faculty) from groups underrepresented in medicine were competitively selected from a national pool of applicants (2016-2019). TRANSCENDS activities comprised an online Clinical Research degree program, monthly webinars, AAN meeting activities, and mentoring. Participants were surveyed during and after completion of TRANSCENDS to evaluate program components.</p> <p class="MsoNoSpacing"><b>Outcomes:</b> Of 23 accepted scholars (comprising four successive cohorts), 56% were women; 61% Hispanic/Latinx, 30% Black/African American, 30% assistant professors. To date, 48% have graduated the TRANSCENDS program and participants have published 180 peer-reviewed articles.   Mentees' feedback noted that professional skills development (i.e., manuscript and grant writing), networking opportunities, and mentoring were the most beneficial elements of the program. Stated opportunities for improvement included: incorporating a mentor-the-mentor workshop, providing more transitional support for mentees in the next stage of their careers, and requiring mentees to provide quarterly reports.</p> <p class="MsoNoSpacing"><b>Conclusions:</b> TRANSCENDS is a feasible program for supporting underrepresented in medicine neurologists towards careers in research and faculty academic appointments attained thus far have been sustained. While longer term outcomes and process enhancements are warranted, programs like this may help increase the numbers of diverse academic neurologists, and further drive neurological innovation.</p>

opencc-zeroApr 2022View details →
zenodo36/100

Revitalizing Middle School Classrooms: Integrating Scholarly Literature and Survey Data to Foster Student Engagement

<p>Literature and survey data. Triangualtion matrices</p>

opencc-by-4.0May 2023View details →
dryad36/100

Data from: TRANSCENDS: a career development program for underrepresented in medicine scholars in academic neurology

Open the record for dataset details and reuse information.

publicApr 2022View details →
zenodo32/100

Social Network Data for Jones, The Decline and Fall of the Assyrian Court Scholar

Open the record for dataset details and reuse information.

opencc-by-4.0Oct 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record