Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
396
datasets available to search
ShareScore release 0.7.1
Dataset results
396 results for “Books”
The glosses to the first book of the Etymologiae of Isidore of Seville: raw data
<p>This excel file contains the raw data behind the digital scholarly edition of the glosses to the first book of the <em>Etymologiae</em> of Isidore of Seville published at: <a href="https://db.innovatingknowledge.nl/edition">https://db.innovatingknowledge.nl/edition</a></p>
Dialogic Density in the Books of William James and John Dewey
<p>This study employs a machine learning algorithm (the Stanford Named Entity Recognizer, or NER) to shed light on the relative rates at which William James and John Dewey mention other persons in their respective books. The NER attempts to tag words and phrases in a corpus with either PERSON, ORGANIZATION, or LOCATION. I created databases for every monograph by James and Dewey. Each database contains all entities tagged with PERSON in the relevant book. </p>
Dialogic Density in the Books of William James and John Dewey
<p>This study employs a machine learning algorithm (the Stanford Named Entity Recognizer, or NER) to shed light on the relative rates at which William James and John Dewey mention other persons in their respective books. The NER attempts to tag words and phrases in a corpus with either PERSON, ORGANIZATION, or LOCATION. I created a corpus of major books published by James and Dewey, respectively, and used the NER to analyze each corpus. I then created databases and collected all entities tagged with PERSON in each book in each corpus. </p>
Fig. 4 in The Amount And Distribution Of The Red Data Book Bird Wetland Species In The Azov-Black Sea Region Of Ukraine According To The Results Of August Counts 2004-2015
Fig. 4. Distribution of wetlands number depending on number of species in them (axis X — number of species, axis Y — number of wetlands).
Fig. 5 in The Amount And Distribution Of The Red Data Book Bird Wetland Species In The Azov-Black Sea Region Of Ukraine According To The Results Of August Counts 2004-2015
Fig. 5. Distribution of wetlands depending on species number (axis X) and average amount of birds in them (axis Y).
Books from the OpenAIRE Research Graph
<p>This dataset is the subset of the OpenAIRE Research Graph about research products of type "Book".</p> <p>The tar archive contains gz files, each with one json per line. Each json compliant to the schema available at <a href="http://doi.org/10.5281/zenodo.5799514">http://doi.org/10.5281/zenodo.5799514</a>. </p>
Computing Book Parts with EEBO-TCP
<p>Full frequency data for the div type attributes used in EEBO-TCP files to accompany the article 'Computing Book Parts with EEBO-TCP'.</p>
YALTAi: Segmonto Manuscript and Early Printed Book Dataset
<p>This dataset has been built to train a segmentation model. It contains ALTO and YOLOv5 formats</p> <p>This dataset is derived from:</p> <ul> <li>CREMMA Medieval ( Pinche, A. (2022). Cremma Medieval (Version Bicerin 1.1.0) [Data set]. https://github.com/HTR-United/cremma-medieval )</li> <li>CREMMA Medieval Lat (Clérice, T. and Vlachou-Efstathiou, M. (2022). Cremma Medieval Latin [Data set]. https://github.com/HTR-United/cremma-medieval-lat )</li> <li>Eutyches. (Vlachou-Efstathiou, M. Voss.Lat.O.41 - Eutyches "de uerbo" glossed [Data set]. https://github.com/malamatenia/Eutyches)</li> <li>Gallicorpora HTR-Incunable-15e-Siecle ( Pinche, A., Gabay, S., Leroy, N., & Christensen, K. Données HTR incunable du 15e siècle [Computer software]. https://github.com/Gallicorpora/HTR-incunable-15e-siecle )</li> <li>Gallicorpora HTR-MSS-15e-Siecle ( Pinche, A., Gabay, S., Leroy, N., & Christensen, K. Données HTR manuscrits du 15e siècle [Computer software]. https://github.com/Gallicorpora/HTR-MSS-15e-Siecle )</li> <li>Gallicorpora HTR-imprime-gothique-16e-siecle ( Pinche, A., Gabay, S., Vlachou-Efstathiou, M., & Christensen, K. HTR-imprime-gothique-16e-siecle [Computer software]. https://github.com/Gallicorpora/HTR-imprime-gothique-16e-siecle )</li> </ul> <p>+ a few hundred newly annotated data, specifically the test set which is completely novel and based on early prints and manuscripts.</p> <p> </p> <table> <tbody> <tr> <td>Dataset</td> <td>Number of images</td> </tr> <tr> <td>Train</td> <td>854</td> </tr> <tr> <td>Dev</td> <td>154</td> </tr> <tr> <td>Test</td> <td>139</td> </tr> </tbody> </table> <p> </p>
Books in a bubble : assessing the OAPEN Library collection
<p>Open access infrastructure for books is becoming more mature, and it is being used by an increasing number of people. The growing importance of open access infrastructure leads to more interest in sustainability, governance and impact assessment. For this paper, we will assess the OAPEN Library. In the spring of 2022, it passed the milestone of 20,000 titles. This was a good moment to evaluate the core asset of the OAPEN Library: its collection.</p> <p>How well meets the collection the needs of its users? The OAPEN Library sees global usage; the collection reflects this by offering titles in over 50 languages. The collection is not focused on a specific subject area, but the choice of medium – books and chapters, not journals and articles – is more strongly associated with the humanities and social sciences. It does not track its users, but the supporters of the OAPEN Libraries are globally distributed academic institutions, scientific and scholarly funders and publishers. An assessment of the OAPEN Library should therefore take into account the diversity of languages, subjects and users.</p>
Illustrations from the Environmental Data Science Book: Shared under CC-BY 4.0 for reuse
<p>Illustrations as part of the <em>Environmental Data Science</em> book.</p> <p>When using any of the images, please include the following attribution with the specific DOI as listed on the particular Zenodo page:</p> <blockquote> <p>This illustration is created by Scriberia with The Turing Way community. Used under a CC-BY 4.0 licence. DOI: <a href="https://doi.org/10.5281/zenodo.7030142">10.5281/zenodo.7030142</a></p> </blockquote> <p>When using any of the images, please include the following attribution with the specific DOI as listed on the particular Zenodo page:</p> <p>You can cite all versions by using the DOI <a href="https://doi.org/10.5281/zenodo.7030142">10.5281/zenodo.7030142</a>. This DOI represents all versions, and will always resolve to the latest one.</p> <p><em>This work was supported by Wave 1 of The UKRI Strategic Priorities Fund under the EPSRC Grant EP/W006022/1, particularly the Environment & Sustainability theme within that grant & The Alan Turing Institute.</em></p>
Multidimensional impact assessment of a large collection of books using PlumX: methodology, technical limitations and indicators analysis [Complementary material to manuscript]
<p>The main purpose of this macro-study is to shed light on the broad impact of books. For this purpose, the impact a very large collection of books (more than 200,000) has been analysed by using PlumX, an analytical tool providing a great number of different metrics provided by various tools. Furthermore, the study focuses on the evolution of the most significant measures and indicators over time. The results show usage counts in comparison to the other metrics are quantitatively predominant. Catalogue holdings and reviews represent a book’s most characteristic measures deriving from its increased level of impact in relation to prior results. Our results also corroborate the long half-life of books within the scope of all metrics, excluding views and social media. Despite of some disadvantages, PlumX has proved to be a very helpful and promising tool for assessing the broad impact of books, especially because of how easy it is to enter the ISBN directly as well as its algorithm to aggregate all the data generated by the different ISBN variations.</p>
Table 1: The Number of Words and Lessons in Book Three
<p>Three hundred and thirty three lexical items included in the English textbook, namely,<br> Birjandy, P. & Norouzi, M. & Mahmoody, G. (2004). English Book Three, which is routinely<br> assigned by the Ministry of Education for EFL teaching in Iranian public high schools in Grade<br> three were taught to the participants in the present study. The book typically includes sections of<br> reading comprehension, grammar, pronunciation practice, dialogues, and a list of new words (with<br> no explanation of meaning or synonyms) in the ending part of each lesson. These wordlists (See<br> Appendix 1) were used for vocabulary instruction and subsequent vocabulary testing in the study.<br> To save space, the number of the new words in each lesson is presented in Table 1</p>
Dataset: Booking Holdings Inc. (BKNG) Stock Performance
This dataset provides historical stock market performance data for specific companies. It enables users to analyze and understand the past trends and fluctuations in stock prices over time. This information can be utilized for various purposes such as investment analysis, financial research, and market trend forecasting.
Books per Publisher: Data from the Directory of Open Access Books (DOAB)
<p>Quantitative information on publishers and published books from the Directory of Open Access Books (DOAB).</p>
Numbers of Articles, Books and Dissertation theses indexed in BASE and percentages of items published Open Access, under Creative Commons Licenses and under Open Licenses (2013-2017)
<p>A look at the data provided by the Open Access search engine BASE (http://base-search.net) shows that the Open Science compliance among dissertation theses stagnates. BASE knows three categories of accessibility: Open Access, Unknown, Non-Open Access. In the following tables and graphs, figures reported as "Open Access" have been categorised by BASE as Open Access. The tables and graphics show data from BASE (as of 06.03.2018) as follows:</p> <p>a) Indexed theses, books and journal articles</p> <p>b) Indexed theses, books and journal articles published by Open Access</p> <p>c) indexed theses, books and journal articles under Creative Commons licenses.</p> <p>d) indexed theses, books and journal articles, which are published under open licenses in the sense of the Open License, i. e. reflect terms of use of the Open Source.</p> <p> </p> <p>Although doctoral theses already had a high share of open access by 2013 (43%), by 2017 it had risen by only 5% (2017:48%). At the same time, the proportion of books published in open access rose by 14% (from 20% to 34%) and articles by 17% from 44% (2013) to 61% (2017). The same effect can be seen in the proportion of CC-licensed items: Their share rose by 4% (from 9% to 13%) for doctoral theses, by 9% for books (from 4% to 13%) and 8% for articles (from 10% to 18%) between 2013 and 2017. However, the share of openly licensed items is most pronounced: it did not increase for doctoral theses, but remained at 2% between 2013 and 2017; in the same period it increased by 5% (from 1% to 6%) for books, and by 5% (from 5% to 10%) for articles.</p>
Medieval Book Hand Scribes
<p>This is a data set comprising 46 medieval scribes writing in book hand scripts. Each scribe is represented by 20 manuscript images (pages or spreads). Documentation in the file doc.pdf.</p>
Data and Code For PhD Thesis- Making an Impression: An Assessment of the Role of Print Surfaces Within the Technological, Commercial, Intellectual and Cultural Trajectory of Book Illustration c.1780-c.1860.
<p><strong>Overview</strong></p> <p>These datasets support research found in the PhD Thesis entitled: Making an Impression: An Assessment of the Role of Print Surfaces Within the Technological, Commercial, Intellectual and Cultural Trajectory of Book Illustration c.1780-c.1860, submitted by William Finley. The data was provided by the British Library as part of a wider ambition to digitise millions of book illustrations (further details can be found here: <a href="https://github.com/BL-Labs/imagedirectory">https://github.com/BL-Labs/imagedirectory</a>). The main dataset contains 6063 rows and 108,784 illustrations. All of the data is held in comma-separated values (csv) files. The collection has been subdivided in order to interrogate the dataset further. All of the datasets have been interrogated using Rscript (R 3.4.2 El Capitan Build), the codes of which have been included in the dataset repository. </p> <p><strong>Guidance to Datasets</strong></p> <p>'Counts of Illustration by Size over Time'</p> <ul> <li>Title: Title of Book</li> <li>Author: Author of Book </li> <li>Pub_Place: Location the book was published</li> <li>Book_Id: British Library image identifier </li> <li>YEAR: Year the book was published</li> <li>Image_Count: Number of illustrations belonging to that book</li> <li>N: Number of illustrations according to the size of illustration</li> <li>IMAGE_TYPE: Size of Illustration </li> </ul> <p>'Density Graphs of the Position of Illustrations on the Page'</p> <ul> <li>X1: The X position in pixels of the top left of the image</li> <li>Y1: The Y position in pixels of the top left of the image</li> <li>X2: Width in pixels of the Image</li> <li>Y2: Height in pixels of the Image</li> <li>PERCENT_PAGE: Percentage of the page taken up by illustrations</li> <li>TYPE: Percentage Range of the Page taken up by illustrations </li> <li>VOL: Volume</li> </ul> <p>'Frequency of Illustrations Across First 100 Pages of the Book 1800-1850'</p> <ul> <li>Page: Page number </li> <li>X1: The X position in pixels of the top left of the image</li> <li>Y1: The Y position in pixels of the top left of the image</li> <li>X2: Width in pixels of the image</li> <li>Y2: Height in pixels of the image</li> <li>PERCENTPAGE: Percentage of the page occupied by illustration</li> <li>SIZE: Size category each illustration belongs to</li> </ul> <p>'Illustration Arrangement in Single Books and Editions'</p> <ul> <li>Book_ID: British Library image identifier</li> <li>Page_NO: Page Number Illustration is found on </li> <li>X: The X position in pixels of the top left of the image</li> <li>Y: The Y position in pixels of the top left of the image </li> <li>WIDTH: Width in pixels of the image</li> <li>HEIGHT: Height in pixels of the image</li> <li>IMAGE_SIZE: Percentage of the page occupied by illustration</li> </ul> <p>'Relative Frequency of Printing Methods Over Time'</p> <ul> <li>Publisher: Publisher of the book</li> <li>Title: Title of the book</li> <li>first_author: Author</li> <li>pub_place: Location the book was published</li> <li>book_identifier: British Library image identifier</li> <li>Year: Year the book was published </li> <li>COUNT_IMAGE: Number of illustrations found in a given book</li> <li>YEAR_BOOK_COUNT: Number of book published in a given year</li> <li>YEAR_IMAGE: Number of Illustrations found in books published in a given year</li> <li>AVERAGE_IMAGE: Average number of illustrations found in books published in a given year</li> <li>IMAGE_METHOD: Print method used to produce the illustration</li> <li>IMAGE_SIZE_COUNT: Number of illustrations printed in books in a given year according to its size</li> <li>IMAGE_SIZE: Categories of image sizes</li> </ul> <p><strong>Contact</strong></p> <p>Will Finley can be contacted via the following email. Information listed is accurate at the time of publication</p> <p>Email: wafinley1@sheffield.ac.uk</p> <p> </p> <p> </p>
LIBER 2019 Workshop. Open Access books in academic libraries – how can we adapt workflows and cost management to an open scholarly communications landscape?
<p>This dataset includes all the results from a workshop held at the LIBER Annual Conference 2019 on June 26, 2019, in Dublin, Ireland. The workshop aimed at collecting and discussing current library practices related to open access books.</p> <p>The dataset includes information from a survey made in preparation for the conference, where 67 European libraries responded to a questionnaire based on activities or workflows in libraries related to open access books. Both the survey questionnaire and the results from the survey are uploaded as separate files.</p> <p>The dataset also includes the presentation made by keynote speaker Eelco Ferwerda from the OAPEN Foundation. He presented results based on the 2017 landscape study report on open access monographs with some new results from more recent studies by Springer Nature and a follow-up report by Knowledge Exchange.</p> <p>Olaf Siegert from ZBW - Leibniz Information Centre for Economics presented a brief overview of what libraries can do to promote OA books in terms of collection management, publication services and development of staff and organisation. The conclusion is that it is not necessarily big changes that are needed.</p> <p>Sofie Wennström presented results from a survey aimed at European research libraries on behalf of the LIBER Open Access Working Group. The survey reveals that many libraries are already working with processes to promote OA books. This is done by libraries organising publishing services or inhouse publishing, by including OA books in discovery services and repositories and by supporting authors to learn more about open access and open licensing.</p> <p>Finally, the LIBER Open Access Working Group shares a report from the workshop providing some quick takeaways and some good examples brought up during the breakout session with the workshop participants.</p>
Dataset of Pages from Early Printed Books with Multiple Font Groups
<p>This dataset is composed of photos of various resolution of 35'623 pages of printed books dating from the 15th to the 18th century. Each page has been attributed by experts from one to five labels corresponding to the font groups used in the text, with two extra-classes for non-textual content and fonts not present in the following list: Antiqua, Bastarda, Fraktur, Gotico Antiqua, Greek, Hebrew, Italic, Rotunda, Schwabacher, and Textura.</p> <p>Note that to make downloading the dataset with slow or unreliable Internet connections easier, the dataset has been separated in several zip files. All zip files must be extracted in the same folder. The CSV files containing the labels should ideally be in the parent folder.</p> <p>The labels are provided in two CSV files, one for training/tuning font group recognition methods, and the second one for evaluation purposes. Where several pages come from the same book, a special care has been taken to have all of them in the same subset.</p> <p>The paper presenting this dataset in detail is "Dataset of Pages from Early Printed Books with Multiple Font Groups", accepted at the 5th International Workshop on Historical Document Imaging and Processing, Sydney, Australia.</p> <p>We would like to thank the British Library (London), Bayerische Staatsbibliothek München, Staatsbibliothek zu Berlin, Universitätsbibliothek Erlangen, Universitätsbibliothek Heidelberg, Staats- und Universitäatsbibliothek Göttingen, Stadt- und Universitätsbibliothek Köln, Württembergische Landesbibliothek Stuttgart and Herzog August Bibliothek Wolfenbüttel for the data they sent us and kindly allowed us to use for this public dataset.</p>
Fig. 14 in An Atlas Of Book Lung Fine Structure In The Order Scorpiones (Arachnida)
Fig. 14. Broteochactas delicatus (Karsch, 1879), 1 ³ (AMNH), KOH-macerated cuticle: dorsal view of attachment site of poststigmaticus muscle on posterior edge of book lung spiracle. Posterior edge normally closes spiracle unless pulled open by contraction of poststigmaticus muscle.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.