Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

39

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

39 results for “Scientific literature”

Learn how ShareScore rates datasets ↗
zenodo48/100

Topic Labels of "Dynamic Topic Modelling for Exploring the Scientific Literature on Coronavirus: An Unsupervised Labelling Technique"

<p>These are the labels generated with the method proposed in the article <em>"Dynamic Topic Modelling for Exploring the Scientific Literature on Coronavirus: An Unsupervised Labelling Technique".</em> These labels are for the 100 and 200 DTM topic models, trained both with the whole corpus and with only the COVID-19 period data&nbsp;</p> <p>&nbsp;</p> <p>For the generation of these labels you can go to the original published work or to the linked Zenodo resource.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Sample Records: Disinformation as a strategy of obstructionism on climate action: analysis of the limitations of the scientific literature for a systemic understanding of the phenomenon

<p>The project contains several underlying datasets essential for replicating the study's findings. The dataset <strong>01.1_PRIMERPRISMA_IDENTIFICATION.xlsx</strong> includes the initial selection of 6 general terms related to environment and sustainability and 11 specific terms related to disinformation, summarizing the selected keywords, generated Boolean operators, and initial search results, yielding 783 records. The <strong>01.2_PRIMER PRIMA-SCREENING.xlsx</strong> file details the screening process, eliminating duplicates and non-English documents, resulting in 271 retained records. The <strong>01.3_PRIMER PRISMA_INCLUDED.xlsx</strong> file contains results after further screening, retaining 82 documents with expanded bibliometric details. The <strong>02.1_SEGUNDOPRISMA_IDENTIFICATION.xlsx</strong> file documents the second phase of identification using new terms related to climate and disinformation, retrieving 174 records. The <strong>02.2_SEGUNDOPRISMA_SCREENING.xlsx</strong> file includes the screening process for the second phase, reducing records to 75, with an abstract review retaining 2 documents. The <strong>02.3_SEGUNDOPRISMA_INCLUDED.xlsx</strong> file integrates documents from both search phases and other sources, culminating in a final review of 86 documents. The <strong>3.1_Other sources.xlsx</strong> file includes additional relevant sources identified during the review process. Finally, the <strong>4-Final included.xlsx</strong> file contains the final set of 75 publications subjected to the DESLOCIS analysis model.</p>

opencc-zeroMay 2024View details →
zenodo44/100

Sample Records (Analytical procedure): Disinformation as a strategy of obstructionism on climate action: analysis of the limitations of the scientific literature for a systemic understanding of the phenomenon

<p>This dataset includes t<span>he online form and the results from the quantitative phase of the study: Disinformation as an obstructionist strategy in climate change mitigation: A review of the scientific literature for a systemic understanding of the phenomenon</span></p> <p>To duplicate the form you can use: https://forms.office.com/Pages/ShareFormPage.aspx?id=6sSEXw03nkuDDHVvi_G1Hw0s3dVrMb1NsO12gDNTB9BUREo4WENRMFFDN1lOSlRSU0xJNkVHWURWUS4u&amp;sharetoken=rg4Qfg19O4UgYzUB084C&nbsp;</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Sample Records (PRISMA Checklist and Flow diagram): Disinformation as a strategy of obstructionism on climate action: analysis of the limitations of the scientific literature for a systemic understanding of the phenomenon

<p>This dataset includes: the PRISMA Checklist and the&nbsp;<span>PRISMA Flow diagram of the study titled: Disinformation as an obstructionist strategy in climate change mitigation: A review of the scientific literature for a systemic understanding of the phenomenon.</span></p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Dataset from the study "Analysis of the accuracy of scientific literature references provided by ChatGPT"

<p>This dataset corresponds to the study carried out to analyse 10 bibliographic references of 10 Spanish authors in the field of Information Sciences requested to the ChatGPT chatbot.</p> <p>The file &quot;Bibliographic_references_ analysis&quot; contains the 10 references returned by ChatGPT for each of the 10 authors (a total of 100 references), together with the variables analysed to check their authenticity.</p> <p>The &quot;Keywords_analysis&quot; file contains the normalisation carried out on the words considered to be key words extracted from the titles of the works, according to which a word cloud showing the frequency of occurrence could be drawn up.</p>

opencc-by-4.0Mar 2023View details →
edi44/100

Pond data: physical, chemical, and biological characteristics with scientific and United States of America state definitions from literature and legislative surveys

Ponds are often identified by their small size and shallow depths, but the lack of a universal definition hampers science and weakens legal protection. In order to determine a working definition of ‘pond’, we conducted a literature search for scientific definitions, a U.S. state survey for management definitions, and looked at pond ecosystem function using data from the literature search. Our dataset includes physical, chemical, and biological data for 1327 waterbodies ≤ 20 ha in surface area and ≤ 9 m in maximum or mean depth from our literature review. These data have a global distribution, we include a table of latitudes and longitudes, and span many years (1946-2019). We have also included a table of 54 pond definitions from the literature review and a table of U.S. state definitions of ponds, wetlands, and lakes resulting from our survey.

openCC (other)Apr 2022View details →
zenodo40/100

Organizing the fragmented landscape of multidisciplinary product development: A mapping of approaches, processes, methods and tools from the scientific literature - Searchable cartographies

<p>This document gathers cartographies for the development of mechatronic products, cyber-physical systems and smart products. The three cartographies presented are associated with an open-access article &ndash; see the citation box below&nbsp;&ndash; and differ from the ones provided in the article in that they are searchable, which makes it easier to pinpoint references, concepts and techniques. This document comprises a legend, the cartographies and a list of associated references.&nbsp;</p> <p>To contextualize the cartographies, the integration of digital and connectivity technologies in new products can invite companies to adapt their development. Organizing the fragmented landscape of multidisciplinary product development to help companies navigate the dense scientific literature corpus is a first step in supporting them in doing so. Multidisciplinary product development can be investigated by analyzing specific types of products that deal with both software and hardware development and can be referred to as cyber-physical systems, mechatronics, and smart products and systems in the literature. To support their development, 236&nbsp;&ldquo;concepts and techniques&rdquo; (an expression that encompasses approaches, processes, methods and tools) were identified from 167&nbsp;scientific papers through an extensive literature review and organized based on a four-level model paired with a decision tree. The mapping of the sorted concepts and techniques made it possible to generate graphical representations called &ldquo;cartographies.&rdquo; These cartographies represent a database of concepts and techniques for multidisciplinary product development and serve to support companies in their transformation from the product development perspective by providing them with a general overview of the related literature.</p>

opencc-by-4.0Feb 2022View details →
zenodo40/100

Supporting Material: Scientific Literature on the Sustainable Development Goals (SDGs). Scopus - May 2022

<p>This is a supplementary dataset for an article analysing the scientific literature related to the Sustainable Development Goals (SDGs) using Scopus-indexed journals.&nbsp;</p> <p>Data were retrieved in May 2022. Scopus was searched in the Title, Abstract, and Keywords fields&nbsp; looking for each of the 17 SDGs (search query example: TITLE-ABS-KEY (&ldquo;SDG1&rdquo; or "SDG 1").</p> <p>The dataset includes the following information for each of the 4808 scientific publication:</p> <p>ID: an identificatory alphanumerical number given by the authors</p> <p>Primary SDG: the main SDG the document focus on (MULTIPLE in case of more than one, ALL in case of all the SDGs)</p> <p>Year: Year of publication</p> <p>Title: title of the publication</p> <p>Abstract: abstract of the publication</p> <p>Index keywords: keywords of the publication</p> <p>&nbsp;</p>

opencc-by-4.0Apr 2024View details →
zenodo40/100

Figure 3. Guettarda speciosa L. A in Katot yan panao: A case study of indigenous botanical nomenclature in the scientific literature

Figure 3. Guettarda speciosa L. A) Common habitus as shrubby tree, Guam (130198175). B) Leaves crowded at branch terminus, Guam (12922175). C) Flowers and buds, Saipan (7100716). D) Fruit, Aitutaki, Cook Islands (163554764). Image numbers from iNaturalist (www.inaturalist.org); photographers: C. Certeza (A), M. Freedman (B), M. Kargul (C), A. Chapman (D); licensing: CC BY-NC (A–C), CC BY-NC-SA (D).

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 4. A in Katot yan panao: A case study of indigenous botanical nomenclature in the scientific literature

Figure 4. A) Type specimens of Claoxylon marianum Muell.-Arg. (G 00313924). B) Gaudichaud's field number " 248 " with Chamoru name " Catud Cunau (Catoud Counao) ". C) Field number " 63 " and the Chamoru name " Panao ". D) A note indicating " Ç'est plutôt un Claxylon[sic]! Juss. " E) Gaudichaud ' s signature and date. F–G) Page six of Gaudichaud's inventory of Mariana plants in which he originally listed specimen 63 as " Guettarda ". Images © Conservatoire et Jardin botaniques de la Ville de Genève with permission (A–E) and public domain, courtesy F. Wamprechts (F–G).

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 2 in Katot yan panao: A case study of indigenous botanical nomenclature in the scientific literature

Figure 2. Claoxylon marianum Muell.Arg. (Euphorbiaceae), Guam. A) Common habitus as shrubby tree (94524278). B) Toothed leaves crowded at branch terminus (86964495). C) Male flowers and buds (153656635). D) Female flowers and fruit (86964496). Image numbers from iNaturalist (www.inaturalist.org); photographers: N. Sablan (A, C), M. Martinez (B, D); licensing: © the author with permission (A, C), CC BY-NC (B, D).

opencc-by-4.0Jun 2024View details →
zenodo40/100

Figure 1 in Katot yan panao: A case study of indigenous botanical nomenclature in the scientific literature

Figure 1. Dendrocnide latifolia (Gaud.) Chew (Urticaceae). A) Common habitus as shrubby tree, Saipan (104853909). B) Leaves crowded at branch terminus, Rota (12946963). C) Female flowers and leaf abscission scars on branches, Guam (34476422). D) Male flowers, Guam (34476420). Image numbers from iNaturalist (www.inaturalist.org); photographers: H. Rogers (A), M. Freedman (B), PACN Vegetation Program (C–D); licensing: © the author with permission (A), CC BY-NC (B–D).

opencc-by-4.0Jun 2024View details →
zenodo40/100

SheepNet scientific literature database

<p>Searchable database containing reviewed and catalogued papers of relevance to the 3 SheepNet themes: reproductive efficiency, gestation efficiency and lamb survival.</p>

opencc-by-4.0Oct 2019View details →
zenodo40/100

Ocean multi-use scientific literature bibliographies

<p>Bibtex files listing and describing scientific publications dealing with ocean multi-use and multiple uses published until the end of 2020. The file called &quot;2O210209_Scopus_final_KWclean.bib&quot; is a collection of publications dealing with multi-use in its broader sense while the file named &quot;20210209_Scopus_final_short_KWclean.bib&quot; is a selection of publications refering to multi-use in its narrower sense (i.e. synergistic combinations of human activities at sea).&nbsp;</p> <p>Both file are the result of the following process:</p> <p>The bibliographic search was performed on Scopus on publications&rsquo; title, abstract and keywords as follows: &ldquo;multi-use&rdquo; OR &ldquo;multiple uses&rdquo; OR &ldquo;multifunctional use&rdquo; OR &ldquo;co-use&rdquo; AND &ldquo;ocean&rdquo; OR &ldquo;sea&rdquo; OR &ldquo;marine&rdquo; OR &ldquo;maritime&rdquo; OR &ldquo;coastal&rdquo;. Filters were used to exclude conference papers, notes and non-classified publications. This query returned 1 700 distinct documents published between 1970 and 2020. After individually reviewing each one, 1 389 publications were purged from the corpus because they were exclusively focused on the terrestrial realm, dealing with other topics than human activities at sea or approaching marine uses separately.</p> <p>Once these steps were completed, the corpus contained 311 references, including 278 journal papers and 33 book chapters. Some papers published before 2010 and others dealing with <em>pescatourism</em> were not captured by Scopus since they did not explicitly refer to MU or its synonyms. In the first case, the authors mentioned in the title, the abstract or the keywords the uses combined instead of multi-use and, in the second one, they do not always label their research with this term. In spite of these limitations, the corpus was larger and more diverse than expected. The papers dealing with the European conception of multi-use were embedded in a large number of publications mainly related to Marine Protected Areas (MPAs) and secondarily to Marine Spatial Planning (MSP). In fact, protected perimeters allowing human activities such as fishing or tourism, as well as marine spaces covered by planning processes are often described as &ldquo;multiple uses&rdquo; and even &ldquo;multi-use&rdquo; territories. Schupp et al. already mentioned that ocean multi-use was linked to the management model inspired by the Great Barrier Reef Marine Park zoning experience, but they did not explain how these two objects of study were connected. This relationship deserves special attention since MU, MPAs and MSP address, albeit quite differently, the same problem: the long-term co-existence of intensifying and diversifying activities at sea. Thus, it was decided to extract the 68 papers related to the European conception of multi-use to compare this sub-corpus to the main one. The first one is referred as the short collection and to the second one as the large collection.</p>

opencc-by-4.0Apr 2023View details →
dryad36/100

Data from: Biocultural approaches to sustainability: a systematic review of the scientific literature

Current sustainability challenges demand approaches that acknowledge a plurality of human-nature interactions and worldviews, for which biocultural approaches are considered appropriate and timely. This systematic review analyses the application of biocultural approaches to sustainability in scientific journal articles published between 1990 and 2018 through a mixed methods approach combining qualitative content analysis and quantitative multivariate methods. The study identifies seven distinct biocultural lenses, i.e. different ways of understanding and applying biocultural approaches, which to different degrees consider the key aspects of sustainability science - inter and transdisciplinarity, social justice and normativity. The review suggests that biocultural approaches in sustainability science need to move from describing how nature and culture are co-produced to co-producing knowledge for sustainability solutions, and in so doing, better account for questions of power, gender and transformations, which has been largely neglected thus far.

opencc-zeroJun 2020View details →
zenodo36/100

The growth of COVID-19 scientific literature: A forecast analysis of different daily time series in specific settings

<p>Submitted to&nbsp;The ISSI 2021 Conference.&nbsp;The conference is organised by KU Leuven in close collaboration with the university of Antwerp under the auspices of ISSI &ndash; the International Society for Informetrics and Scientometrics (<a href="http://www.issi-society.org/">http://www.issi-society.org/</a>).&nbsp;</p> <p>We present a forecasting analysis on the growth of scientific literature related to COVID-19 expected for 2021. Considering the paramount scientific and financial efforts made by the research community to find solutions to end the COVID-19 pandemic, an unprecedented volume of scientific outputs is being produced. This questions the capacity of scientists, politicians and citizens to maintain infrastructure, digest content and take scientifically informed decisions. A crucial aspect is to make predictions to prepare for such a large corpus of scientific literature. Here we base our predictions on the ARIMA model and use two different data sources: the Dimensions and World Health Organization COVID-19 databases. These two sources have the particularity of including in the metadata information on the date in which papers were indexed.&nbsp; We present global predictions, plus predictions in three specific settings: by type of access (Open Access), by NLM source (PubMed and PMC), and by domain-specific repository (SSRN and MedRxiv). We conclude by discussing our findings.</p>

opencc-by-4.0Jan 2021View details →
zenodo36/100

Identification of focus versus background entities in scientific literature

<p>This dataset allows training and evaluating methods for the identification of focus versus background entities in scientific literature. A focus entity is an entity being actively research in a publication while a background entity is an entity that is being discussed in a publication but is not the main focus of the publication.</p> <p>The dataset has been generated automatically using the MeSH indexing of MEDLINE as reference. The entities of interest in this dataset are microbial pathogens. Entities were annotated using a dictionary approach and then the MeSH indexing of the MEDLINE citation linked to the publication was used to determine the relevance of the entity as focus or background entity.</p> <p>There are two main types of datasets, one generated from MEDLINE (files medline.*) and another one generated using full text articles from PubMed Central articles (PMC) (files pmc.*). The data sets are split into training and test, which we used in our research. All fields within the files are separated using the pipe &quot;|&quot; character. The MEDLINE citation dataset contains data from over 1M citations while the PMC dataset from over 100k publications (which is a subset of the MEDLINE dataset). In each row in the dataset files, the pathogen of interest has been replaced by the text @PATHOGEN$ and there might be several references of the pathogen in the same row.</p> <p>Full text articles datasets have been further split into a dataset with explicit separation between sections and another one in which all the full text article appears in one single text string and section names appear at the beginning of each section.</p>

opencc-by-4.0Jan 2022View details →
zenodo36/100

Dataset for "Unleashing the Power of Knowledge Extraction from Scientific Literature in Catalysis"

<p>JCIM paper link: <a href="https://pubs.acs.org/doi/10.1021/acs.jcim.2c00359">https://pubs.acs.org/doi/10.1021/acs.jcim.2c00359</a><br><br>Github repo: <a href="https://github.com/nsndimt/CatalysisIE">https://github.com/nsndimt/CatalysisIE</a><br><br><br>Dataset Content:</p> <ul> <li>Pretrained BERT:&nbsp;<code>scibert_domain_adaption.tar.gz</code>&nbsp;extract it to&nbsp;<em>pretrained</em> directory</li> <li>Cross-Validation Checkpoint:&nbsp;<code>cross_validation_checkpoint.tar.gz</code>&nbsp;extract it to&nbsp;<em>checkpoint</em> directory</li> <li>Annotated Data: <code>data.jsonl</code> and <code>split.jsonl</code>&nbsp;put it under&nbsp;<em>data</em> directory</li> </ul>

opencc-by-4.0May 2022View details →
zenodo36/100

LLM-Based Knowledge Graph Construction from Materials Research Scientific Literature

<p>This dataset was constructed by creating a benchmark of 349 manually annotated triples, which were extracted from four different research articles in the field of materials science.</p>

opencc-by-4.0Oct 2024View details →
dryad36/100

Data from: Biocultural approaches to sustainability: a systematic review of the scientific literature

Open the record for dataset details and reuse information.

publicJun 2020View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record