Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
49
datasets available to search
ShareScore release 0.9.0
Dataset results
49 results for “bibliographical data”
Bibliographic Data from the Digital Twin Anomaly Detection Decision-Making for Bridge Management Systematic Review
<p>This database contains all the bibliographic information about the 8673 records found after applying the Search Strategy used for the Digital Twin Anomaly Detection Decision-Making for Bridge Management Systematic Review. Such strategy consisted on using seven initial keywords and similar terms of interest (namely: bridge and bridges, etc.): </p> <ul> <li>Bridge.</li> <li>Digital twin.</li> <li>Bridge information modelling.</li> <li>Finite elements.</li> <li>Bridge health monitoring.</li> <li>Anomaly detection algorithm.</li> <li>Cultural heritage.</li> </ul> <p>Six initial queries were done combining the first keyword with the rest of them:</p> <ul> <li>bridge* AND "digital twin*"</li> <li>bridge* AND (BrIM OR "bridge information model*")</li> <li>bridge* AND (FEM OR FEA OR "finite element method*" OR "finite element analy*")</li> <li>bridge* AND ("bridge health monitoring" OR "structural health monitoring")</li> <li>bridge* AND (ADA OR "anomaly detection algorithm*")</li> <li>bridge* AND ("cultural heritage" OR "monument* bridge*" OR "old bridge*" OR "ancient bridge*" OR "historic* bridge*")</li> </ul> <p>As a first screening step, the combination of these 6 initial searches was done to obtain relevant works containing at least three of the main keywords of interest:</p> <ul> <li>#1 AND #2</li> <li>#1 AND #3</li> <li>#1 AND #4</li> <li>#1 AND #5</li> <li>#1 AND #6</li> <li>#2 AND #3</li> <li>#2 AND #4</li> <li>#2 AND #5</li> <li>#2 AND #6</li> <li>#3 AND #4</li> <li>#3 AND #5</li> <li>#3 AND #6</li> <li>#4 AND #5</li> <li>#4 AND #6</li> <li>#5 AND #6</li> </ul> <p>All records found in Scopus where downloaded both in .ris and .csv format and are included in this database. The search was conducted on 10/12/2022.</p> <p>Note: Searches 10, 14, 17 and 21 did not return any records.</p>
Bibliographic Data from the Computational Methods Applied to Earthen Historical Structures Review
<p>This database contains all the bibliographic information about the 293 records found after applying the Search Strategy used for the Computational Methods Applied to Earthen Historical Structures Review. Such strategy consisted on using relevant keywords grouped into three different search queries within ”TITLE-ABS-KEY”, for the years 2019-2023:</p> <ol> <li>(”earthen heritage” OR ”earthen historical building*” OR ”earthen historical structure*” OR ”earthen architect*” OR ”earthen monument*”).</li> <li>(adobe OR ”rammed earth” OR cob ) AND (”computational method*” OR ”numerical analy*”).</li> <li>(adobe OR ”rammed earth” OR cob ) AND (fem OR dem OR la OR ”finite element” OR ”discrete element” OR ”limit analysis”).</li> </ol> <p>The search was conducted on April 7, 2023.</p>
Bibliographic Data from the SoTL in Civil and Structural Engineering Systematic Review
<p>This database contains all the bibliographic information found after applying the Search Strategy used for the SoTL in Civil and Structural Engineering Systematic Review. The following electronic databases were searched:</p> <ul> <li>Scopus.</li> <li>Web of Science.</li> <li>OsloMet Library.</li> <li>Google Scholar (no bibliographic information is presented since this database does not allow to download such data).</li> </ul> <p>A total of 84 records were found in Scopus, 43 in Web of Science, and 55 in OsloMet Library. The search was conducted on September 1, 2023.</p> <p>The information is presented in .ris, .bib, and .csv format.</p>
Supplementary material 3: World Spider Catalog Bibliographic Data: Treatments from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063
List of journal/publisher by ranked by treatment count exported from the World Spider Catalog 14 October 2014 with total treatments by source, cumulative treatments, and cumulative proportion of treatments.
Supplementary material 2: World Spider Catalog Bibliographic Data: Publications from: Integrating and visualizing primary data from prospective and legacy taxonomic literature - Biodiversity Data Journal 3: e5063 (12 May 2015) https://doi.org/10.3897/BDJ.3.e5063
Ranked list of journal/publisher exported from the World Spider Catalog 14 October 2014 with total articles by source, cumulative articles, and cucmulative proportion of articles.
Raw and aggregated data for the study introduced in the paper "The way we cite: common metadata used across disciplines for defining bibliographic references"
<p>These data have been gathered in the context of a study aiming to investigate citation practices for referencing different types of entities and, in particular, for understanding the most used metadata in bibliographic references. The data are stored in two documents in XLSX format:</p> <ul> <li>file "links-intext-pointers-and-cited-entity-types.xlsx" - it contains information about whether the in-text reference pointers of the various PDF articles of the corpus have specified hypertextual links from the in-text reference pointers to the denoted bibliographic reference, plus information about the types of all the entities cited by each article in the corpus;</li> <li>file "metadata-bibliographic-references.xlsm" - it contains information about the metadata used to identify the various descriptive elements of all the bibliographic references defined in the article of the corpus.</li> </ul> <p>The methodology used to gather all these data is described in:</p> <blockquote> <p>Santos, E. A. d., Peroni, S., Mucheroni, M. L.: Workflow for retrieving all the data of the analysis introduced in the article "Citing and referencing habits in Medicine and Social Sciences journals in 2019". (2020), <a href="https://doi.org/10.17504/protocols.io.bbifikbn">https://doi.org/10.17504/protocols.io.bbifikbn</a></p> </blockquote>
Global suicide mortality rates (2000-2019) and bibliographic data
<p>The dataset contains World Bank Suicide mortality rate WDI (world development indicator) (2000-2019) world-wide data in original and processed form. In addition to the statistical data this dataset also contains bibliographic records of articles published on the topic of suicide in relation to individual countries during (2000-2019) in original and processed form. </p> <p>The data consists of six archives:</p> <ol> <li>World development indicator suicide mortality rate SH.STA.SUIC.P5. This archive contains suicide mortality rate of 159 countries during the period of 2000-2019 per 100,000 population including males and females as of November, 2023.</li> <li>Web of science records country and suicide. This archive contains bibliographic records organized by country on the topic of suicide related to that country published during 2000-2019 as of November, 2023.</li> <li>Suicide mortality rate statistics and keywords. This archive contains processed data of 1 and 2 archives in three files. The 'Countries suicide rates and WOS records' contains organized temporal suicide mortality rate data for each country and each year for males and females including counts of articles on suicide related in that country. The 'words and countries matrix' file contains information about how many times author and paper keywords from suicide related publications were seen in articles associated with each country. This data is organized as matrix in which rows are keywords, columns are countries and cells are counts of the keyword. The 'words and countries pairs' file contains same information only organized as keyword country pairs.</li> <li>Suicide mortality rate clusters countries keywords titles. This archive contains bibliographic data organized by country clusters. These clusters group countries with similar suicide mortality rate dynamics in males and females shown in two included figures. Each folder of the cluster contains a section with bibliographic records; a section with keywords associated with each country; and a section in which each publication associated with the country has a separate filecontaining its title and keywords.</li> <li>Suicide keywords embedding data. This archive contains word embedding vectors and metadata learned by recurrent neural network trained to classify countries from suicide related keywords of articles associated with those countries. Folder 'trained with keywords' contains embeddings learned in classifying countries in which training samples are keyword strings of publications. Folder 'trained with titles' contains embeddings learned in classifying countries in which training samples are strings containing titles of publication plus keywords.</li> <li>Suicide keywords association rule mining. This archive contains files of subsets of keywords frequently mentioned together in suicide related publications. Folder 'Mining in clusters' has frequent keyword itemsets in country clusters. Folder 'Mining in individual countries' has frequent keyword itemsets in countries. Examples of keyword networks connecting clusters and networks connecting countries in individual clusters are included which helps to identify specific and shared keywords by country clusters and by countries in the individual clusters. </li> </ol> <p>These datasets support a data availability statements for upcoming articles.</p>
Bibliographic dataset based on Scientometrics, containing provenance information compliant with the OpenCitations Data Model and non disambigued authors
<p>The dataset contains bibliographical information about scholarly works in the journal Scientometrics only if the DOI is known. The data was extracted via Crossref. It is a temporal dataset in which provenance information and change-tracking have been managed by adopting the OpenCitations Data Model. Moreover, the dataset contains information on all the cited academic works. Journals and bibliographic resources always appear unambiguously, without duplicates. On the contrary, the authors have not been disambigued. Finally, heuristics have been applied to recover the DOI of the cited works in case Crossref did not provide such information.</p>
Raw and aggregated data for the study introduced in the article "An analysis of citing and referencing habits across all scholarly disciplines: approaches and trends in bibliographic metadata errors"
<p>This dataset contains all the raw data and aggregated data subject of the study introduced in the article "An analysis of citing and referencing habits across all scholarly disciplines: approaches and trends in bibliographic metadata errors". The study is based on the bibliographic and citation data contained in 729 articles published in 147 journals in 27 subject areas. The articles contained a total amount of 34,140 bibliographic references and 55,100 mentions and quotations overall.</p> <p>The dataset is composed of a series of files:</p> <ul> <li>the files "subject_area_<discipline-name>.csv" contain the raw data of the articles published in the journals of all the disciplines considered in the study;</li> <li>the file "article_data_summary.csv" contains the aggregated data created considering the raw data in the previous files, which have been used to creating all the tables and figures in the article;</li> <li>the file "starred_metadata_set.csv" contains information about the most used subset of bibliographic metadata;</li> <li>the file "journals_selection.csv" contains information about all the journals selected for the study.</li> </ul>
Bibliographic dataset based on Scientometrics, including provenance information compliant with the OpenCitations Data Model
<p>The dataset contains bibliographical information about scholarly works in the journal Scientometrics only if the DOI is known. The data was extracted via Crossref. It is a temporal dataset in which provenance information and change-tracking have been managed by adopting the OpenCitations Data Model. Moreover, the dataset contains information on all the cited academic works. Journals, bibliographic resources, and authors always appear unambiguously, without duplicates. Finally, heuristics have been applied to recover the DOI of the cited works in case Crossref did not provide such information.</p> <p>The dataset is distributed as two journal files, one for the data and one for the provenance, readable via the triplestore Blazegraph. There are 4,960,087 data triples and 19,348,027 provenance triples, which corresponds to 1,134,545 entities and 2,696,689 snapshots. Therefore, on average, each entity has two snapshots. Among the data, there are 231,217 agent roles, 221,602 responsible agents, 206,003 bibliographic resources, 142,472 citations, 141,555 bibliographical references, 108,112 identifiers, and 83,584 resource embodiments.</p> <p>The code to generate and modify such collections is available at <a href="https://doi.org/10.5281/zenodo.5579754">https://doi.org/10.5281/zenodo.5579754</a>. </p>
Bibliographic data for the systematic review on tilting table tests of masonry assemblies
<p>This database contains all the bibliographic information about the records found after applying the Search Strategy used for the Systematic Review on Tilting Table Tests of Masonry Assemblies. The search was conducted in Scopus, Web of Science, IEEE Explore, Engineering Village, and Wiley Online Library databases. It was performed on 13/09/2023. The bibliographic data of the records found in the different databases is presented in .ris, .bib, and .csv format.</p>
Bibliographic data and analysis of COVID-19 research outputs from Imperial College London 16.01.2020-02.04.2020
<p>Bibliographic data and analysis of 41 research outputs, including reports/preprints/published articles/code, identified as having Imperial authorship and being relevant to COVID-19, published between 16.01.2020 - 02.04.2020. </p> <p>Related report can be found at: Price RC and Ozkan YA. 13 weeks in a pandemic: a descriptive study of Imperial College London’s COVID-19 publications. Imperial College London (April 2020), https://doi.org/10.25561/77970</p>
Bibliographic data of La trasmissione della conoscenza registrata. Scritti in onore di Mauro Guerrini offerti dagli allievi
<p>The file is in format .ris and contains all the bibliographic references of the festschritf "La trasmissione della conoscenza registrata. Scritti in onore di Mauro Guerrini offerti dagli allievi", edited by Carlo Bianchini and Lucia Sardo, Milano, Editrice Bibliografica, 2021, ISBN: 978-88-9357-347-4</p>
Bibliographic data on datasets affiliated to Poznan University of Technology and indexed in Data Citation Index (retrieved by Web of Science service in January 2023))
<p>The file contains the number of datasets published by the researchers affiliated to Poznan University of Technology and indexed in Data Citation Index provided by Web of Science (database updated 10.01.2023). The Search was performed using the name of institution in the 'Affiliation' field. Dataset contains two files in two diffrent formats: plain text and xls.</p>
Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic
<p>This data set contains supplementary material for the paper 'Large-scale comparison of bibliographic data sources: Scopus, Web of Science, Dimensions, Crossref, and Microsoft Academic' by Martijn Visser, Nees Jan van Eck, and Ludo Waltman. The data set provides the statistics presented in the figures in the paper.</p>
Data for "Modular Bibliographical Profiling of Historic Book Reviews"
<p>This dataset supports the research paper, ""Modular Bibliographical Profiling of Historic Book Reviews." The paper examines different methods of predicting bibliographical details (e.g. author, title, and publisher) of books under review in a corpus of approximately 1,100 historical book reviews. The dataset is comprised of book reviews from ProQuest's American Periodicals Series (APS). This kind of bibliographical profiling is often characterized as a Natural Language Processing (NLP) or Named Entity Recognition (NER) task, but it can be more specifically described as a two-part Named Entity Linking (NEL) task, beginning with a feature extraction stage followed by one of several available matching or classification methods. An attempt has been made to formalize constraints for modular bibliographical profiling (MBP) and shed light on some important choices that are often glossed over or obscured by digital humanities (DH) practitioners. Applying these constraints, the paper evaluates combinations of feature selection (naive bag-of-words [BOW], rule-based feature extraction, and NER using a pre-trained model) with a standardized similarity-based matching strategy (cosine similarity). All tasks are performed on derived text data (term frequency tables), so that data can be shared and all methods can be used on materials available only in non-consumptive formats. These comparisons suggest that naive BOW can perform quite robustly, and that using even a basic pretrained NER model in conjunction with a BOW approach may reduce false positives. </p>
An Analysis of the Current Bibliographical Data Landscape in the Humanities. A Case for the Joint Bibliodata Agendas of Public Stakeholders - video presentation
<p>A video presenting the DARIAH's Bibliographical Data Working Group entitled <em>An Analysis of the Current Bibliographical Data Landscape in the Humanities. A Case for the Joint Bibliodata Agendas of Public Stakeholders. </em>The report is freely available on Zenodo: <a href="https://zenodo.org/record/6559857#.Y0XDo3ZBy5f">https://zenodo.org/record/6559857#.Y0XDo3ZBy5f</a>. </p> <p>This presentation aims to present the original work - co-authored by 18 WG's members - in a condensed manner.</p>
Databases and information systems for research output: digital humanities outlook (DARIAH Bibliographical Data Working Group, September, 30th 2022)
<p>The video recording of the workshop "Databases and information systems for research output: digital humanities outlook" organized by DARIAH Bibliographical Data Working Group.</p>
PubMed inner references obtained from five freely available bibliographic data sources
<p>This dataset contains PMID-to-PMID citations of PubMed 2020 Baseline extracted from five freely available bibliographic data sources (COCI, Dimensions, MAG, NIH-OCC, and S2ORC).</p> <p>Each line contains one citing PubMed document and its cited references. The citing and cited documents are separated by a tab (\t) and the cited references are separated by a semicolon (;).</p>
Bibliographic data of "Mostra Medici in guerra", University of Milano-Bicocca
<p>The file is in format .ris and contains all the bibliographic references of the exhibition “Medici in Guerra. Testimonianze del primo conflitto mondiale dagli archivi storici della Bicocca”, organized by the University of Milano-Bicocca (14th March 2023 – 12nd May 2023).</p> <p>The exhibition pertains the Italian doctors, neurologists, psychiatrists and psychologists who participated in the World War I in their youth. A tragic experience that influenced their lives, leaving a mark on their professional paths as well.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.