Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

19

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

19 results for “OpenAlex”

Learn how ShareScore rates datasets ↗
zenodo44/100

Experimental AI corpus from OpenAlex

<p>A corpus of AI research from OpenAlex. Includes:</p> <ul> <li>A works table with metadata about AI papers</li> <li>An authors&nbsp;table with information about the authors</li> <li>An institutions table with information about institutions</li> <li>A concepts table with information about concepts in works</li> <li>A MeSH table with information about MeSH terms in works</li> <li>A concepts json with the OpenAlex concept taxonomy</li> <li>An abstracts json with deinverted abstracts</li> <li>A citations json with citations from papers</li> </ul> <p>See `ai_openalex_description.md` for data dictionaries.</p> <p>See `ai_openalex_methodology.md` for a description of the method used to create the dataset.</p> <p>See here for additional information:&nbsp;<a href="https://github.com/nestauk/ai_genomics">https://github.com/nestauk/ai_genomics</a></p>

opencc-by-4.0Aug 2022View details →
zenodo44/100

OpenAire Research Graph linked with OpenAlex

<p>This package contains linked datasets of OpenAire Research Graph and OpenAlex.&nbsp;</p> <p>Files descriptions:</p> <p>- author_to_publication_dic.json contains a mapping of authors to their publications</p> <p>- downloads_views_dic.json contains mappings of the publication id to the number of its downloads and views</p> <p>- id_doi_dic.json contains a mapping of the publication id to its doi</p> <p>- merged1..5.json contain all publication data from the OARG dataset</p> <p>- necessary_fields_dic.json contains extracted publications&rsquo; fields necessary for the work</p> <p>- oarg_ref_rel_dic.json contains mapping of publication id to referenced and related work present in OpenAlex dataset</p> <p>- openalex_found_publications5_4.json contains all data on found publications from the OpenAlex</p> <p>- publication_to_author_dic.json contains a mapping of publications to their authors</p>

opencc-by-4.0Aug 2024View details →
zenodo40/100

Datos de instituciones panameñas con ROR en OpenAlex - 2023

<p>Datos de instituciones paname&ntilde;as que aparecen en la plataformas OpenAlex que se identificaron con ROR en el 2023.</p> <p>Diccionario de datos:</p> <ul> <li><strong>Instituci&oacute;n:&nbsp;</strong> nombre de la instituci&oacute;n</li> <li><strong>tipo :</strong> tipo de instituci&oacute;n ( Education, Government, Nonprofit, Facility) (Dato cualitativo)</li> <li><strong>ROR: </strong>ID de ROR de la instituci&oacute;n</li> <li><strong>urlROR: </strong>url de la instituci&oacute;n en ROR</li> <li><strong>works_count;</strong> n&uacute;mero de documentos de la instituci&oacute;n integrados en OpenAlex (datos cuantitativo)</li> </ul>

opencc-by-4.0May 2024View details →
zenodo40/100

OpenAlex Topic Classification v1 Model Artifacts and Training Data

<p>This is all data used to train the topic classification model and also the model artifacts to deploy the model. Please see the github repo for more information:</p> <p>https://github.com/ourresearch/openalex-topic-classification</p>

opencc-zeroJan 2024View details →
zenodo40/100

Enriched OpenAlex Data for Universidad de Antioquia

<p>Advanced user API output from http://impactu.colav.co for the Institutional&nbsp;Profile &quot;Universidad de Antioquia&quot;</p>

opencc-by-4.0Mar 2023View details →
zenodo40/100

OpenAlex Author Name Disambiguation V3 Data - Disambiguation Model

<p>5 Separate files used in the OpenAlex (https://openalex.org) V3 Author Name Disambiguation Model Creation:</p> <ol> <li>ORCID_hard_negative_pairs: Pairs of ORCIDs where either the full name, family name, or given name are a match and would therefore be more difficult to disambiguate.</li> <li>Disambiguator_all_possible_training_data: Dataset created which contains all possible features for modeling and all possible samples of data. Eventually, this was split into train/val/test and also processed more to create a better balance of positive to negative samples for our purposes.</li> <li>Disambiguator_final_train_data: Final data which the disambiguator was trained on.</li> <li>Disambiguator_final_val_data: Data which was used to test the model during training to optimize the features/hyperparameters chosen.</li> <li>Disambiguator_final_test_data: Final dataset which gave model performance indication after all hyperparameters were tuned and features were chosen.</li> </ol> <p>More details can be found at&nbsp;https://github.com/ourresearch/openalex-name-disambiguation</p>

opencc-by-4.0Jul 2023View details →
zenodo36/100

DATASET OF "Exploratory study of OpenAlex datasets to assess productivity of authors and institutions"

<p><strong>Dataset to "Exploratory study of OpenAlex datasets to assess productivity of authors and institutions"</strong></p>

opencc-by-4.0Jul 2024View details →
zenodo36/100

Core sources and core publications in OpenAlex

<p>This data set contains data on core sources and core publications identified in the OpenAlex database (based on the OpenAlex snapshot released on August 30, 2024).</p> <p>The source code used to identify core sources and core publications in OpenAlex is available in <a href="https://github.com/CWTSLeiden/CWTS-OpenAlex-databases/tree/2024aug" target="_blank" rel="noopener">this GitHub repository</a>.</p> <p>See&nbsp;<a href="https://doi.org/10.5281/zenodo.13879947" target="_blank" rel="noopener">this report</a> for more information about the identification of core sources and core publications in OpenAlex.</p> <p>&nbsp;</p> <p>This data set consists of the following tab-delimited files.</p> <p>&nbsp;</p> <p>source.tsv</p> <ul> <li>source_id</li> <li>source</li> <li>source_type</li> <li>issn_l</li> <li>is_core_source</li> <li>n_works</li> <li>n_core_works</li> </ul> <p>&nbsp;</p> <p>&nbsp;work.tsv</p> <ul> <li>work_id</li> <li>work_type</li> <li>pub_year</li> <li>source_id</li> <li>doi</li> <li>is_core_work</li> </ul>

opencc-zeroApr 2024View details →
zenodo36/100

OpenAlex Author Name Disambiguation V3 Initial Clusters

<p>Author name disambiguation V3 initial clusters for the OpenAlex dataset. See <a href="https://openalex.org">https://openalex.org</a></p> <p>There are 633803287 rows, split into 4 CSV (comma-delimited) files (with headers).</p> <p>The CSV files have two columns: &quot;work_author_id&quot; and &quot;author_id&quot;</p> <p>&quot;work_author_id&quot;: An OpenAlex Work ID and an author sequence number, joined with an underscore (&quot;_&quot;)</p> <p>&quot;author_id&quot;: An OpenAlex Author ID, representing a unique author in OpenAlex</p>

opencc-by-4.0Jul 2023View details →
zenodo32/100

More open abstracts? Comparing abstract coverage in Crossref and OpenAlex [dataset]

<p>Aggregated data underlying the blogpost:<strong><br><br>More open abstracts? Comparing abstract coverage in Crossref and OpenAlex<br></strong><a href="https://bmkramer.github.io/SesameOpenScience_site/thought/202411_open_abstracts/">https://bmkramer.github.io/SesameOpenScience_site/thought/202411_open_abstracts/</a><strong><br></strong><br>The dataset contains the following files:</p> <ul> <li><em>abstracts_crossref_openalex_202410.csv</em></li> <li><em>abstracts_crossref_openalex_data_dictionary.txt</em></li> </ul> <p>The csv file contains data on abstract coverage for Crossref DOIs in Crossref and OpenAlex, aggregated by publisher, for the top 1000 publishers in terms of number of retrieved dois. Scope is limited to publications with Crossref type 'journal-articles' and publication years 2022-2024. Variables are described in the data dictionary included in this record.<br><br>This analysis was performed using&nbsp;<a href="https://openknowledge.community/" rel="nofollow">Curtin Open Knowledge Initiative (COKI)</a>&nbsp;infrastructure, which is documented on GitHub:&nbsp;<a href="https://github.com/The-Academic-Observatory">https://github.com/The-Academic-Observatory</a>. Here, a number of open data sources (including Crossref, OpenAlex and OpenAIRE) are ingested into a Google Big Query environment, which can then be queried via SQL.<br><br>The following data sources were used:</p> <ul> <li> <p>Crossref (Metadata Plus snaphot 2024-10-31, Crossref member route API 2024-11-20)</p> </li> <li> <p>OpenAlex (data snapshot 2024-10-30)</p> </li> </ul> <p><br>The code used to generate the dataset is available on GitHub: <a href="https://github.com/bmkramer/more_open_abstracts">https://github.com/bmkramer/more_open_abstracts</a></p> <p>&nbsp;</p>

opencc-zeroNov 2024View details →
zenodo32/100

CSET scholarly literature metadata over OpenAlex works

<p>This dataset contains metadata developed at the Center for Security and Emerging Technology that augments OpenAlex works, including outputs from CSET-developed classifiers. Detailed documentation is available <a href="https://eto.tech/dataset-docs/emerging-technology-overlay-openalex/">here</a>.</p> <p>The attached zip file contains a set of JSONL files which comprise our dataset. Each row conforms to <a href="https://github.com/georgetown-cset/cset_openalex/blob/main/schemas/metadata.json" target="_blank" rel="noopener">this schema</a>, with null values omitted. This dataset is currently a work in progress and full documentation will be made available at a later date.</p> <p>Research subject classifications are based on work supported in part by the Alfred P. Sloan Foundation under Grant No. G-2023-22358.</p>

opencc-by-nc-4.0Apr 2024View details →
zenodo32/100

Enriched OpenAlex Data for Colombia

<p>See&nbsp;https://github.com/colav-playground/advanced_user_tests</p>

opencc-by-4.0Apr 2023View details →
zenodo32/100

Colombia Coauthorship networks divided by openalex lvl 0 concepts

<p>Graphs are uploaded in gml&nbsp;format, which can be easily imported by networkx,&nbsp;gephi and neo4j.</p> <p>The undirected graphs are constructed from openalex database (last update Dic 2022), where nodes are authors and edges specify whether or not two&nbsp;nodes coauthored at least one paper.&nbsp;We avoided papers with more than 10 authors since they are very scarse and could affect the posterior analysis of the networks.</p> <p>The attributes of the nodes consist of the list of used words for each author and its frequency and all the concepts (of every level) the papers of the author are labeld.</p> <p>The attributes of the edges only contains the number of papers published between two authors.</p> <p>For more information about openalex concepts, visit&nbsp;https://docs.openalex.org/api-entities/concepts</p>

opencc-by-4.0Feb 2023View details →
zenodo32/100

OpenAlex Authors and Affiliations, V2

<p>This is the old, deprecated author data for <a href="https://openalex.org">OpenAlex</a>. In July, 2023, the OpenAlex dataset switched to a new author name disambiguation (V3), and deprecated all old author IDs. This is a data dump of those old author IDs.</p> <p>- Author metadata are included in `authors/**/*.gz` files.</p> <p>- Affiliations---mapping of OpenAlex work IDs to (old) Author IDs---are in the `affiliations_export_20230719T1259/*.gz` files.</p> <p>The files are all gzipped JSON-lines files.</p> <p>See the <a href="https://docs.openalex.org">OpenAlex documentation</a> for more information about OpenAlex and how you can use these files.</p>

opencc-by-4.0Jul 2023View details →
zenodo28/100

OpenAlex Book Publications 2013-2022

Open the record for dataset details and reuse information.

opencc-zeroNov 2023View details →
zenodo28/100

2024 dataset on independent researchers collected from OpenAlex

<p>This dataset belongs to a paper about independent researchers submitted for the STI conference 2024 (https://sti2024.org/). It consists of several files described below. The data is from OpenAlex, collected through the InSySPo instance of the february snapshot of OpenAlex, hosted on Google Cloud. Since Topics are a new feature of OpenAlex data and therefore not part of the snapshot, this data as well as some other data not available at the InSySPo instance at the time of collection have been collected through the OpenAlex API, and incorporated in the files. Data from Scopus and Web of Science may be retrieved by using the search string in the appendix of the article.</p> <p><strong>Files all domains</strong></p> <p><em>240307_open_alex_works.tsv</em></p> <p>contains all works retrieved with the search string for Independent researchers in OpenAlex in the article's appendix.</p> <p><strong>Files Social Sciences and/or Arts &amp; Humanities</strong></p> <p><em>240312_open_alex_works_soc_sci_arts_2010.tsv</em></p> <p>contains articles by Independent researchers in Social Sciences and Humanities published from 2010 and retrieved from OpenAlex.</p> <p><em>240312_open_alex_authors_soc_sci_arts_2010.tsv</em></p> <p>contains authors who are Independent researchers in Social Sciences and Humanities published from 2010 and retrieved from OpenAlex.</p> <p><em>240313_open_alex_authors_all_works_soc_sci_arts_2010.tsv</em></p> <p>contains all works by Independent researchers in Social Sciences and Humanities published from 2010 and retrieved from OpenAlex. All works mean that the researcher has at least once indicated independent status in the affiliation, and the author's other works are also included.</p> <p><em>author_distribution_domain1.csv</em></p> <p>contains number of works per number of authors in the domain Social Sciences (includes Arts &amp; Humanities).</p> <p><em>author_distribution_field33.csv</em></p> <p>contains number of works per number of authors in the field Social Sciences.</p> <p><em>author_distribution_field12.csv</em></p> <p>contains number of works per number of authors in the field Arts &amp; Humanities.</p> <p><em>all_ssh_oa.csv</em></p> <p>contains data for analyzing open access patterns for the domain Social Sciences (includes Arts &amp; Humanities).</p>

openApr 2024View details →
zenodo28/100

OpenAlex Snapshot

<div> <div> <p>OpenAlex is an open, comprehensive index of scolarly papers, citations, authors, institutions, and journals. Available through API and UI as well (at openalex.org), this record refers to the full data snapshot.</p> </div> </div> <div>&nbsp;</div> <p><strong>When citing OpenAlex, don't use this record. Instead, use:</strong></p> <p><em>Priem, J., Piwowar, H., &amp; Orr, R. (2022). OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. ArXiv. https://arxiv.org/abs/2205.01833</em></p> <p>This record is intended for long-term persistence but because the OpenAlex snapshot updates every month, it is better to download the current version directly from AWS. Information on how to download the entire data snapshot for OpenAlex can be found at: <a href="https://docs.openalex.org/download-all-data/openalex-snapshot">https://docs.openalex.org/download-all-data/openalex-snapshot</a></p> <p>&nbsp;</p>

opencc-zeroOct 2024View details →
zenodo28/100

OpenAlex slices for "Collaboration and topic switches in Science"

<p>OpenAlex slices stored as zipped parquet files. Needs pandas &gt;= 2, pyarrow &gt;= 7.</p>

opencc-by-4.0Apr 2023View details →
zenodo20/100

Works of Colombian publisher taken from OpenAlex

Open the record for dataset details and reuse information.

opencc-by-4.0Nov 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record