Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,307
datasets available to search
ShareScore release 0.7.1
Dataset results
1,307 results for “libraries”
Public reference library for edible dormouse calls
<p>The acoustics of small mammals, particularly dormice species, are generally understudied. We explored the vocalisations and acoustic behaviour of the edible dormouse (<em>Glis glis</em>) in various environments in Catalonia (northern Iberian Peninsula) using ultrasound recorders between 2022 and 2023. Up to five different types of calls were identified in various environmental conditions (captivity and free-ranging animals) and developmental stages, from pups to adults. Additionally, one new call type was described, highlighting the plasticity of their vocalisations, and ultrasonic sound production was discovered in pups, suggesting ontogenetic changes in the vocal repertoire. With this project, we emphasise the potential of the acoustic method as a non-invasive tool for studying ecological behaviours and interactions, or early detection of the species in the natural environment.</p> <p>We, hereby, provide an open reference library for edible dormouse calls, laying the groundwork for a better understanding of its acoustics, behaviour and conservation. The compilation provides five clean sequences of high quality for each of the five call types recorded under different environmental conditions (totalling 35 recordings). This includes the chirp (captivity and wild), blow (captivity), aggressive call (captivity and wild), pups during manipulation (wild) and free-ranging pups (wild). The collection of sounds was prepared under the Dormouse Project (<a href="http://www.dormice.org/" target="_blank" rel="noopener">www.dormice.org</a>)</p> <p>This work was supported by the Barcelona Zoo Fundation under the Research and Conservation Program Grant; Generalitat de Catalunya under Grant number ARD264/23/000001; and Diputació de Barcelona under Grant number 2023/0005732.</p>
The Pan-Canadian Chemical Library: A Mechanism to Open Academic Chemistry to High-Throughput Virtual Screening
<h1>Pan-Canadian Chemical Library</h1> <p>This Zenodo repository contains the cheap and druglike subset of the Pan-Canadian Chemical Library (PCCL) project. For more information, visit <a href="https://pccl.thesgc.org/" rel="nofollow">https://pccl.thesgc.org</a>.</p> <h2>PCCL library</h2> <p>The PCCL library is splitted by reaction, then by number of heavy atoms. Two types of files are available in zip archives:</p> <ul> <li>The SMILES format files, with the SMILES string and their product name,</li> <li>The CSV format file, with all the information generated during their enumeration: reagents, druglike properties, etc.</li> </ul> <p>Note: Purchasability is defined according to two integers: 1 for products only composed of BB-50 reagents, and 2 for products composed of BB-40 or BB-50 reagents. Read more about the meaning of these reagents groups in the article below.</p> <h2>Citation</h2> <p>If you find the PCCL useful or if you use it, please cite our paper:</p> <p>Bedart, C. <em>et al.</em> The Pan-Canadian Chemical Library: A mechanism to open academic chemistry to high-throughput virtual screening. Scientific Data 11, (2024).<br>doi: <a title="10.1038/s41597-024-03443-5" href="https://www.nature.com/articles/s41597-024-03443-5">10.1038/s41597-024-03443-5</a></p> <p> </p> <p> </p>
DEI in libraries
Relatively recently, a colleague (Peggy Griesinger) distributed a bibliography on the topic of diversity, equity, and inclusion (DEI), and I decided to spend some time analyzing the content of the bibliography. This missive outlines what I was able to extract, given the limited time I spent.
Reading Information Technology and Libraries (volume 41, number 2, June 2022)
Today, for a good time, I applied my Reader Toolbox to the latest issue of ITAL for the two-fold purposes of: 1) just seing whether the Toolbox could function, and 2) determine the degree I could extract meaningful themes from the issue. Well, the Toolbox functioned, in that it did not crash nor output invalid data, and I do believe I could pull out themes, in that each issue's authors wrote about something distinctive, and I could identify those things. Below describes my process.
Reading Journal of eScience Librarianship: Responsible AI in Libraries and Archives
A special issue of Journal of eScience Librarianship was brought to my attention. The issue was on the topic of responsible AI in libraries and archives. I did a bit of distant reading against the issue, and outlined here are some of my take-aways. In short, AI is something to consider in Library Land, but not without some forethought.
FIG. 4 in The Raymond Benoist microslide library of woods of French Guiana at the Herbarium of Paris (P): restoration and comments
FIG. 4. — Xylem details: A, B, Roucheria calophylla Planch. (synonym of Hebepetalum humiriifolium (Planch.) Benth. & Hook.f. ex B.D. Jacks., det. by D. Sabatier 1996, Benoist 1582 [P04755674]) transverse sections; C, D, Vouacapoua americana Aubl. (Mélinon 596 23 [P00793775]), transverse and tangential sections; E, F, Aspidosperma album (Vahl) Benoist ex Pichon (Benoist 331 [P00402200]). Scale bars: A, C, E, 200 µm; B, 50 µm; D, F, 100 µm.
FIG. 2 in The Raymond Benoist microslide library of woods of French Guiana at the Herbarium of Paris (P): restoration and comments
FIG. 2. — Restoration process at the lab: A, restoration battery of histological slides; B, dismantling process of mounting: 1, slide in cold water; 2, removal of section, label and coverslip after 24-48 h at 60°C; C, dehydration and remounting of specimens under the extractor hood.
FIG. 5 in The Raymond Benoist microslide library of woods of French Guiana at the Herbarium of Paris (P): restoration and comments
FIG. 5. — Histological details: xylem: A, B, Lecythis poiteaui Berg (Benoist 192 [P04543005]), transverse and tangential sections; C, D, Unonopsis rufescens (Baill.) R.E. Fries (Benoist 1267 [P02132803]), transverse sections; bark: E, F, Enterolobium schomburgkii (Benth.) Benth. (Benoist 133 [P00199488]), phelloderm and phellem. Scale bars: A, E, F, 100 µm; B, C, 200 µm; D, 50 µm.
FIG. 1 in The Raymond Benoist microslide library of woods of French Guiana at the Herbarium of Paris (P): restoration and comments
FIG. 1. — Raymond Benoist (1881-1970): A, portrait of c. 1930, Muséum national d'Histoire naturelle ©; B, a tray with the arrangement of slides of wood until 2021.
Beyond the Shelves: Embarking on Citizen Science with Your Library
<p>The video “Beyond the Shelves: Embarking on Citizen Science with your Library” by LibOCS project partner UT Library shows how to make first connections with librarians and get to know the role of university libraries in citizen science.</p>
Dataset for "The Ithildin library for efficient numerical solution of anisotropic reaction-diffusion problems in excitable media"
<p>This archive contains the full source code of Ithildin as well as the data generated by the simulations used in the paper introducing the Ithildin software.</p>
Figure 5. A front view of the "Mary and John Gray Library"-Modeling, Designing, and Implementing an Avatar-based Interactive Map
<p>Figure 5 represents the avatar standing outside and in front of the Mary and John Gray Library after selecting the option “Library”. The library’s main purpose is to facilitate students with a variety of scholarly information within the overall composition of the University’s stated mission. Figure 5 shows the path generated by A* algorithms with a red color.</p>
Figure 6. An inside view of Mary and John Gray Library-Modeling, Designing, and Implementing an Avatar-based Interactive Map
<p>Figure 6 exhibits the ambience of the study environment that allows students to have group discussions, and when to access Internet, and more. The photographs have been digitized in a very realistic way.</p>
Hathi Trust Library Vectorized features
<p>A smaller-resolution (and therefore more portable) version of the Stable Random Projection Hathi Trust features described in my forthcoming article. The Northeastern repository is many individual files with 1280 random dimensions; this is just 640 random dimensions. The numbers are also experimentally encoded as half-precision floats, which cuts the file size by half at the cost of only being supported by my Python module. The net result is a file 1/4 the size of the full resolution ones for the paper that has, probably, something like 60-80% of the information content.</p> <p>The full file is '<a href="https://www.zenodo.org/api/files/6d615dbd-65de-4391-93ac-91b302bb57e4/ht-640d-complete-half-precision.bin?versionId=0e9f5551-0888-454c-8bb8-0dd3f3d6d949">ht-640d-complete-half-precision.bin</a>'. You can also download 11 smaller files organized by language.</p> <p>Since these files use half-precision float encoding, to read them you must specify the precision when reading: e.g.,<br> </p> <pre><code class="language-python">from SRP import Vector_file f = Vector_file("ita.bin", precision = "half")</code></pre> <p>Code to read these files is at https://github.com/bmschmidt/pySRP. </p>
Digital Scholarship services and supports - an overview from Irish Research and National Libraries - Data with Comments
<p>This dataset contains survey data with comments (cleaned, direct references to institutions removed) and broken down by survey question sections. Some of the data has been converted to counts where it was used to generate the charts.</p> <p>The data is the output of a 2018 CONUL (Ireland’s consortium of research and national libraries) survey focusing on Digital Scholarship services and supports from a Irish Research and National Library perspective.</p>
Maven 99 most popular library statical usages
<p>A SQL database containing the static usages of API elements of any version of the 99 most used maven artifact, by any of it client on maven central.</p>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 4 of 4)
<p><strong>This is part 4 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 4 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 3 of 4)
<p><strong>This is part 3 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 3 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 2 of 4)
<p><strong>This is part 2 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre> <p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><strong>This is part 2 of 4. The following dataset contains a reference the other ZIP files: </strong></p> <pre>https://doi.org/10.5281/zenodo.2598101</pre>
A biodiversity dataset graph: Biodiversity Heritage Library (BHL)
<p>A biodiversity dataset graph: Biodiversity Heritage Library </p> <p>Biodiversity datasets, or descriptions of biodiversity datasets, are increasingly available through open digital data infrastructures such as the Biodiversity Heritage Library (BHL, https://biodiversitylibrary.org). "The Biodiversity Heritage Library improves research methodology by collaboratively making biodiversity literature openly available to the world as part of a global biodiversity community." - https://biodiversitylibrary.org , June 2019. </p> <p>However, little is known about how these networks, and the data accessed through them, change over time. This dataset provide snapshots of all OCR item texts (e.g., individual items) available through BHL as tracked by Preston (https://github.com/bio-guoda/preston , https://doi.org/10.5281/zenodo.1410543 ) over period May - June 2019.</p> <p>This snapshot contains about 120GB of uncompressed OCR texts across 227k OCR BHL items. Also, a snapshot of the BHL item catalog at https://www.biodiversitylibrary.org/data/item.txt is included.</p> <p>The archive consists of 256 individual parts (e.g., preston-00.tar.gz, preston-01.tar.gz, ...) to allow for parallel file downloads. The archive contains three types of files: index files, provenance files and data files. Only two index and provenance files are included and have been individually included in this dataset publication. Index files provide a way to links provenance files in time to eestablish a versioning mechanism. Provenance files describe how, when and where the BHL OCR text items were retrieved. For more information, please visit https://preston.guoda.bio or https://doi.org/10.5281/zenodo.1410543). </p> <p>To retrieve and verify the downloaded BHL biodiversity dataset graph, first concatenate all the downloaded preston-*.tar.gz files (e.g., cat preston-*.tar.gz > preston.tar.gz). Then, extract the archives into a "data" folder. After that, verify the index of the archive by reproducing the following result:</p> <p>$ java -jar preston.jar history<br> <0659a54f-b713-4f86-a917-5be166a14110> <http://purl.org/pav/hasVersion> <hash://sha256/89926f33157c0ef057b6de73f6c8be0060353887b47db251bfd28222f2fd801a> .<br> <hash://sha256/41b19aa9456fc709de1d09d7a59c87253bc1f86b68289024b7320cef78b3e3a4> <http://purl.org/pav/previousVersion> <hash://sha256/89926f33157c0ef057b6de73f6c8be0060353887b47db251bfd28222f2fd801a> .</p> <p>To check the integrity of the extracted archive, confirm that each line produce by the command "preston verify" produces lines as shown below, with each line including "CONTENT_PRESENT_VALID_HASH". Depending on hardware capacity, this may take a while.</p> <p>$ java -jar preston.jar verify<br> hash://sha256/e0c131ebf6ad2dce71ab9a10aa116dcedb219ae4539f9e5bf0e57b84f51f22ca file:/home/preston/preston-bhl/data/e0/c1/e0c131ebf6ad2dce71ab9a10aa116dcedb219ae4539f9e5bf0e57b84f51f22ca OK CONTENT_PRESENT_VALID_HASH 49458087<br> hash://sha256/1a57e55a780b86cff38697cf1b857751ab7b389973d35113564fe5a9a58d6a99 file:/home/preston/preston-bhl/data/1a/57/1a57e55a780b86cff38697cf1b857751ab7b389973d35113564fe5a9a58d6a99 OK CONTENT_PRESENT_VALID_HASH 25745<br> hash://sha256/85efeb84c1b9f5f45c7a106dd1b5de43a31b3248a211675441ff584a7154b61c file:/home/preston/preston-bhl/data/85/ef/85efeb84c1b9f5f45c7a106dd1b5de43a31b3248a211675441ff584a7154b61c OK CONTENT_PRESENT_VALID_HASH 519892</p> <p>Note that a copy of the java program "preston", preston.jar, is included in this publication. The program runs on java 8+ virtual machine using "java -jar preston.jar", or in short "preston". </p> <p>Files in this data publication:</p> <p>README - this file</p> <p>preston-[00-ff].tar.gz - preston archives containing BHL OCR item texts, their provenance and a provenance index.</p> <p>9e8c86243df39dd4fe82a3f814710eccf73aa9291d050415408e346fa2b09e70 - preston index file<br> 2a5de79372318317a382ea9a2cef069780b852b01210ef59e06b640a3539cb5a - preston index file</p> <p>89926f33157c0ef057b6de73f6c8be0060353887b47db251bfd28222f2fd801a - preston provenance file<br> 41b19aa9456fc709de1d09d7a59c87253bc1f86b68289024b7320cef78b3e3a4 - preston provenance file<br> </p> <p>This work is funded in part by grant NSF OAC 1839201 from the National Science Foundation.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.