Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
19
datasets available to search
ShareScore release 0.9.0
Dataset results
19 results for “OpenAlex”
Experimental AI corpus from OpenAlex
<p>A corpus of AI research from OpenAlex. Includes:</p> <ul> <li>A works table with metadata about AI papers</li> <li>An authors table with information about the authors</li> <li>An institutions table with information about institutions</li> <li>A concepts table with information about concepts in works</li> <li>A MeSH table with information about MeSH terms in works</li> <li>A concepts json with the OpenAlex concept taxonomy</li> <li>An abstracts json with deinverted abstracts</li> <li>A citations json with citations from papers</li> </ul> <p>See `ai_openalex_description.md` for data dictionaries.</p> <p>See `ai_openalex_methodology.md` for a description of the method used to create the dataset.</p> <p>See here for additional information: <a href="https://github.com/nestauk/ai_genomics">https://github.com/nestauk/ai_genomics</a></p>
OpenAire Research Graph linked with OpenAlex
<p>This package contains linked datasets of OpenAire Research Graph and OpenAlex. </p> <p>Files descriptions:</p> <p>- author_to_publication_dic.json contains a mapping of authors to their publications</p> <p>- downloads_views_dic.json contains mappings of the publication id to the number of its downloads and views</p> <p>- id_doi_dic.json contains a mapping of the publication id to its doi</p> <p>- merged1..5.json contain all publication data from the OARG dataset</p> <p>- necessary_fields_dic.json contains extracted publications’ fields necessary for the work</p> <p>- oarg_ref_rel_dic.json contains mapping of publication id to referenced and related work present in OpenAlex dataset</p> <p>- openalex_found_publications5_4.json contains all data on found publications from the OpenAlex</p> <p>- publication_to_author_dic.json contains a mapping of publications to their authors</p>
Datos de instituciones panameñas con ROR en OpenAlex - 2023
<p>Datos de instituciones panameñas que aparecen en la plataformas OpenAlex que se identificaron con ROR en el 2023.</p> <p>Diccionario de datos:</p> <ul> <li><strong>Institución: </strong> nombre de la institución</li> <li><strong>tipo :</strong> tipo de institución ( Education, Government, Nonprofit, Facility) (Dato cualitativo)</li> <li><strong>ROR: </strong>ID de ROR de la institución</li> <li><strong>urlROR: </strong>url de la institución en ROR</li> <li><strong>works_count;</strong> número de documentos de la institución integrados en OpenAlex (datos cuantitativo)</li> </ul>
OpenAlex Topic Classification v1 Model Artifacts and Training Data
<p>This is all data used to train the topic classification model and also the model artifacts to deploy the model. Please see the github repo for more information:</p> <p>https://github.com/ourresearch/openalex-topic-classification</p>
Enriched OpenAlex Data for Universidad de Antioquia
<p>Advanced user API output from http://impactu.colav.co for the Institutional Profile "Universidad de Antioquia"</p>
OpenAlex Author Name Disambiguation V3 Data - Disambiguation Model
<p>5 Separate files used in the OpenAlex (https://openalex.org) V3 Author Name Disambiguation Model Creation:</p> <ol> <li>ORCID_hard_negative_pairs: Pairs of ORCIDs where either the full name, family name, or given name are a match and would therefore be more difficult to disambiguate.</li> <li>Disambiguator_all_possible_training_data: Dataset created which contains all possible features for modeling and all possible samples of data. Eventually, this was split into train/val/test and also processed more to create a better balance of positive to negative samples for our purposes.</li> <li>Disambiguator_final_train_data: Final data which the disambiguator was trained on.</li> <li>Disambiguator_final_val_data: Data which was used to test the model during training to optimize the features/hyperparameters chosen.</li> <li>Disambiguator_final_test_data: Final dataset which gave model performance indication after all hyperparameters were tuned and features were chosen.</li> </ol> <p>More details can be found at https://github.com/ourresearch/openalex-name-disambiguation</p>
DATASET OF "Exploratory study of OpenAlex datasets to assess productivity of authors and institutions"
<p><strong>Dataset to "Exploratory study of OpenAlex datasets to assess productivity of authors and institutions"</strong></p>
Core sources and core publications in OpenAlex
<p>This data set contains data on core sources and core publications identified in the OpenAlex database (based on the OpenAlex snapshot released on August 30, 2024).</p> <p>The source code used to identify core sources and core publications in OpenAlex is available in <a href="https://github.com/CWTSLeiden/CWTS-OpenAlex-databases/tree/2024aug" target="_blank" rel="noopener">this GitHub repository</a>.</p> <p>See <a href="https://doi.org/10.5281/zenodo.13879947" target="_blank" rel="noopener">this report</a> for more information about the identification of core sources and core publications in OpenAlex.</p> <p> </p> <p>This data set consists of the following tab-delimited files.</p> <p> </p> <p>source.tsv</p> <ul> <li>source_id</li> <li>source</li> <li>source_type</li> <li>issn_l</li> <li>is_core_source</li> <li>n_works</li> <li>n_core_works</li> </ul> <p> </p> <p> work.tsv</p> <ul> <li>work_id</li> <li>work_type</li> <li>pub_year</li> <li>source_id</li> <li>doi</li> <li>is_core_work</li> </ul>
OpenAlex Author Name Disambiguation V3 Initial Clusters
<p>Author name disambiguation V3 initial clusters for the OpenAlex dataset. See <a href="https://openalex.org">https://openalex.org</a></p> <p>There are 633803287 rows, split into 4 CSV (comma-delimited) files (with headers).</p> <p>The CSV files have two columns: "work_author_id" and "author_id"</p> <p>"work_author_id": An OpenAlex Work ID and an author sequence number, joined with an underscore ("_")</p> <p>"author_id": An OpenAlex Author ID, representing a unique author in OpenAlex</p>
More open abstracts? Comparing abstract coverage in Crossref and OpenAlex [dataset]
<p>Aggregated data underlying the blogpost:<strong><br><br>More open abstracts? Comparing abstract coverage in Crossref and OpenAlex<br></strong><a href="https://bmkramer.github.io/SesameOpenScience_site/thought/202411_open_abstracts/">https://bmkramer.github.io/SesameOpenScience_site/thought/202411_open_abstracts/</a><strong><br></strong><br>The dataset contains the following files:</p> <ul> <li><em>abstracts_crossref_openalex_202410.csv</em></li> <li><em>abstracts_crossref_openalex_data_dictionary.txt</em></li> </ul> <p>The csv file contains data on abstract coverage for Crossref DOIs in Crossref and OpenAlex, aggregated by publisher, for the top 1000 publishers in terms of number of retrieved dois. Scope is limited to publications with Crossref type 'journal-articles' and publication years 2022-2024. Variables are described in the data dictionary included in this record.<br><br>This analysis was performed using <a href="https://openknowledge.community/" rel="nofollow">Curtin Open Knowledge Initiative (COKI)</a> infrastructure, which is documented on GitHub: <a href="https://github.com/The-Academic-Observatory">https://github.com/The-Academic-Observatory</a>. Here, a number of open data sources (including Crossref, OpenAlex and OpenAIRE) are ingested into a Google Big Query environment, which can then be queried via SQL.<br><br>The following data sources were used:</p> <ul> <li> <p>Crossref (Metadata Plus snaphot 2024-10-31, Crossref member route API 2024-11-20)</p> </li> <li> <p>OpenAlex (data snapshot 2024-10-30)</p> </li> </ul> <p><br>The code used to generate the dataset is available on GitHub: <a href="https://github.com/bmkramer/more_open_abstracts">https://github.com/bmkramer/more_open_abstracts</a></p> <p> </p>
CSET scholarly literature metadata over OpenAlex works
<p>This dataset contains metadata developed at the Center for Security and Emerging Technology that augments OpenAlex works, including outputs from CSET-developed classifiers. Detailed documentation is available <a href="https://eto.tech/dataset-docs/emerging-technology-overlay-openalex/">here</a>.</p> <p>The attached zip file contains a set of JSONL files which comprise our dataset. Each row conforms to <a href="https://github.com/georgetown-cset/cset_openalex/blob/main/schemas/metadata.json" target="_blank" rel="noopener">this schema</a>, with null values omitted. This dataset is currently a work in progress and full documentation will be made available at a later date.</p> <p>Research subject classifications are based on work supported in part by the Alfred P. Sloan Foundation under Grant No. G-2023-22358.</p>
Enriched OpenAlex Data for Colombia
<p>See https://github.com/colav-playground/advanced_user_tests</p>
Colombia Coauthorship networks divided by openalex lvl 0 concepts
<p>Graphs are uploaded in gml format, which can be easily imported by networkx, gephi and neo4j.</p> <p>The undirected graphs are constructed from openalex database (last update Dic 2022), where nodes are authors and edges specify whether or not two nodes coauthored at least one paper. We avoided papers with more than 10 authors since they are very scarse and could affect the posterior analysis of the networks.</p> <p>The attributes of the nodes consist of the list of used words for each author and its frequency and all the concepts (of every level) the papers of the author are labeld.</p> <p>The attributes of the edges only contains the number of papers published between two authors.</p> <p>For more information about openalex concepts, visit https://docs.openalex.org/api-entities/concepts</p>
OpenAlex Authors and Affiliations, V2
<p>This is the old, deprecated author data for <a href="https://openalex.org">OpenAlex</a>. In July, 2023, the OpenAlex dataset switched to a new author name disambiguation (V3), and deprecated all old author IDs. This is a data dump of those old author IDs.</p> <p>- Author metadata are included in `authors/**/*.gz` files.</p> <p>- Affiliations---mapping of OpenAlex work IDs to (old) Author IDs---are in the `affiliations_export_20230719T1259/*.gz` files.</p> <p>The files are all gzipped JSON-lines files.</p> <p>See the <a href="https://docs.openalex.org">OpenAlex documentation</a> for more information about OpenAlex and how you can use these files.</p>
OpenAlex Book Publications 2013-2022
Open the record for dataset details and reuse information.
2024 dataset on independent researchers collected from OpenAlex
<p>This dataset belongs to a paper about independent researchers submitted for the STI conference 2024 (https://sti2024.org/). It consists of several files described below. The data is from OpenAlex, collected through the InSySPo instance of the february snapshot of OpenAlex, hosted on Google Cloud. Since Topics are a new feature of OpenAlex data and therefore not part of the snapshot, this data as well as some other data not available at the InSySPo instance at the time of collection have been collected through the OpenAlex API, and incorporated in the files. Data from Scopus and Web of Science may be retrieved by using the search string in the appendix of the article.</p> <p><strong>Files all domains</strong></p> <p><em>240307_open_alex_works.tsv</em></p> <p>contains all works retrieved with the search string for Independent researchers in OpenAlex in the article's appendix.</p> <p><strong>Files Social Sciences and/or Arts & Humanities</strong></p> <p><em>240312_open_alex_works_soc_sci_arts_2010.tsv</em></p> <p>contains articles by Independent researchers in Social Sciences and Humanities published from 2010 and retrieved from OpenAlex.</p> <p><em>240312_open_alex_authors_soc_sci_arts_2010.tsv</em></p> <p>contains authors who are Independent researchers in Social Sciences and Humanities published from 2010 and retrieved from OpenAlex.</p> <p><em>240313_open_alex_authors_all_works_soc_sci_arts_2010.tsv</em></p> <p>contains all works by Independent researchers in Social Sciences and Humanities published from 2010 and retrieved from OpenAlex. All works mean that the researcher has at least once indicated independent status in the affiliation, and the author's other works are also included.</p> <p><em>author_distribution_domain1.csv</em></p> <p>contains number of works per number of authors in the domain Social Sciences (includes Arts & Humanities).</p> <p><em>author_distribution_field33.csv</em></p> <p>contains number of works per number of authors in the field Social Sciences.</p> <p><em>author_distribution_field12.csv</em></p> <p>contains number of works per number of authors in the field Arts & Humanities.</p> <p><em>all_ssh_oa.csv</em></p> <p>contains data for analyzing open access patterns for the domain Social Sciences (includes Arts & Humanities).</p>
OpenAlex Snapshot
<div> <div> <p>OpenAlex is an open, comprehensive index of scolarly papers, citations, authors, institutions, and journals. Available through API and UI as well (at openalex.org), this record refers to the full data snapshot.</p> </div> </div> <div> </div> <p><strong>When citing OpenAlex, don't use this record. Instead, use:</strong></p> <p><em>Priem, J., Piwowar, H., & Orr, R. (2022). OpenAlex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. ArXiv. https://arxiv.org/abs/2205.01833</em></p> <p>This record is intended for long-term persistence but because the OpenAlex snapshot updates every month, it is better to download the current version directly from AWS. Information on how to download the entire data snapshot for OpenAlex can be found at: <a href="https://docs.openalex.org/download-all-data/openalex-snapshot">https://docs.openalex.org/download-all-data/openalex-snapshot</a></p> <p> </p>
OpenAlex slices for "Collaboration and topic switches in Science"
<p>OpenAlex slices stored as zipped parquet files. Needs pandas >= 2, pyarrow >= 7.</p>
Works of Colombian publisher taken from OpenAlex
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.