Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
1,307
datasets available to search
ShareScore release 0.7.1
Dataset results
1,307 results for “libraries”
Developer Expertise Dataset on JavaScript Libraries
<p>This dataset contains an anonymized list of surveyed developers who provided their expertise level on three popular JavaScript libraries:</p> <ol> <li><a href="https://github.com/facebook/react">ReactJS</a>, a library for building enriched web interfaces </li> <li><a href="https://github.com/mongodb/node-mongodb-native">MongoDB</a>, a driver for accessing MongoDB databased </li> <li><a href="https://github.com/socketio/socket.io">Socket.IO</a>, a library for realtime communication </li> </ol>
GALAKSIENN: the G.A.S semi-analytical model associtated library
<p>The GALAKSIENN library contains data sets generated by the G.A.S. semi-analytical model of galaxy formation and evolution. The G.A.S. model is fully described in a set of three Astronomy & Astrophysics papers: G.A.S. I: A prescription for turbulence-regulated star formation and its impact on galaxy properties; G.A.S. II: Dust extinction in galaxies, luminosity functions and InfraRed Excess and G.A.S. III: The panchromatic view of galaxies, Stellar/dust continua and main gas emission lines. The library stores mock galaxy catalogues and ascii tables (stellar mass functions, luminosity functions, number counts ...).</p>
Metadata, Title Pages, and Network Graph of the Digitized Content of the Berlin State Library (146,000 items)
<p>The data set has been downloaded via the OAI-PMH endpoint of the Berlin State Library/Staatsbibliothek zu Berlin’s Digitized Collections (<a href="https://digital.staatsbibliothek-berlin.de/oai">https://digital.staatsbibliothek-berlin.de/oai</a>) on March 1<sup>st</sup> 2019 and converted into common tabular formats on the basis of the provided Dublin Core metadata. It contains 146,000 records.</p> <p>In addition to the bibliographic metadata, representative images of the works have been downloaded, resized to a 512 pixel maximum thumbnail image and saved in JPEG format. The image data is split into title pages and first pages. Title pages have been derived from structural metadata created by scan operators and librarians. If this information was not available, first pages of the media have been downloaded. In case of multi-volume media, title pages are not available.</p> <p>In total, 141,206 images title/first pages are available.</p> <p> </p> <p>Furthermore, the tabular data has been cleaned and extended with geo-spatial coordinates provided by the OpenStreetMap project (<a href="https://www.openstreetmap.org">https://www.openstreetmap.org</a>). The actual data processing steps are summarized in the next section. For the sake of transparency and reproducibility, the original data taken from the OAI-PMH endpoint is still present in the table.</p> <p> </p> <p>To conclude with, various graphs in GML file format are available that can be loaded directly into graph analysis tools such as Gephi (<a href="https://gephi.org/">https://gephi.org/</a>).</p> <p> </p> <p>The implementation of the data processing steps (incl. graph creation) are available as a Jupyter notebook provided at <a href="https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb">https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb</a>.</p> <p> </p> <p>Tabular Metadata</p> <p> </p> <p>The metadata is available in Excel (cleanedData.xlsx) and CSV (cleanedData.csv) file formats with equal content.</p> <p>The table contains the following columns. Italique columns have not been processed.</p> <p>· <em>title</em> The title of the medium</p> <p>· <em>creator</em> Its creator (family name, first name)</p> <p>· <em>subject</em> A collection’s name as provided by the library</p> <p>· <em>type</em> The type of medium</p> <p>· <em>format</em> A MIME type for full metadata download</p> <p>· <em>identifier</em> An additional identifier (most often the PPN)</p> <p>· <em>language</em> A 3-letter language code of the medium</p> <p>· <em>date</em> The date of creation/publication or a time span</p> <p>· <em>relation</em> A relation to a project or collection a medium has been digitized for.</p> <p>· <em>coverage</em> The location of publication or origin (ranging from cities to continents)</p> <p>· <em>publisher</em> The publisher of the medium.</p> <p>· <em>rights</em> Copyright information.</p> <p>· <em>PPN</em> The unique identifier that can be used to find more information about the current medium in all information systems of Berlin State Library/Staatsbibliothek zu Berlin.</p> <p>· spatialClean In case of multiple entries in coverage, only the first place of origin has been extracted. Additionally, characters such as question marks, brackets, or the like have been removed. The entries have been normalized regarding whitespaces and writing variants with the help of regular expressions.</p> <p>· dateClean As the original date may contain various format variants to indicate unclear creation dates (e.g., time spans or question marks), this field contains a mapping to a certain point in time.</p> <p>· spatialCluster The cluster ID determined with the help of the Jaro-Winkler distance on the spatialClean string. This step is needed because the spatialClean fields still contain a huge amount of orthographic variants and latinizations of geographic names.</p> <p>· spatialClusterName A verbal cluster name (controlled manually).</p> <p>· latitude The latitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>· longitude The longitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>· century A century derived from the date.</p> <p>· textCluster A text cluster ID on the basis of a k-means clustering relying on the title field with a vocabulary size of 125,000 using the tf*idf model and k=5,000.</p> <p>· creatorCluster A text cluster ID based on the creator field with k=20,000.</p> <p>· titleImage The path to the first/title page relative to the img/ subdirectory or None in case of a multi-volume work.</p> <p>Other Data</p> <p> </p> <p><em>graphs.zip</em></p> <p> </p> <p>Various pre-computed graphs.</p> <p><em> </em></p> <p><em>img.zip</em></p> <p> </p> <p>First and title pages in JPEG format.</p> <p> </p> <p><em>json.zip</em></p> <p> </p> <p>JSON files for each record in the following format:</p> <p> </p> <p>ppn "PPN57346250X"</p> <p>dateClean "1625"</p> <p>title "M. Georgii Gutkii, Gymnasii Berlinensis Rectoris Habitus Primorum Principiorum, Seu Intelligentia; Annexae Sunt Appendicis loco Disputationes super eodem habitu tum in Academia Wittebergensi, tum in Gymnasio Berlinensi ventilatae"</p> <p>creator "Gutke, Georg"</p> <p>spatialClusterName "Berlin"</p> <p>spatialClean "Berolini"</p> <p>spatialRaw "Berolini"</p> <p>mediatype "monograph"</p> <p>subject "Historische Drucke"</p> <p>publisher "Kallius"</p> <p>lat "52.5170365"</p> <p>lng "13.3888599"</p> <p>textCluster "45"</p> <p>creatorCluster "5040"</p> <p>titleImage "titlepages/PPN57346250X.jpg"</p>
Extract from the Berlin State Library's Main Catalog
<p>The data set is based on the main catalog of the library. Currently, the following fields are extracted:</p> <ul> <li>title</li> <li>author (+ optional GND ID)</li> <li>publisher</li> <li>place of publication</li> <li>country of publication</li> <li>year of publication</li> </ul> <p>The extract has been created by the processPicaPlus script available <a href="https://github.com/elektrobohemian/StabiHacks">here</a>. Attention, some special characters might not have been extracted correctly in versions <1.0.0.</p> <p><em>Change Log:</em></p> <p>0.2.0 </p> <p>fixes various encoding issues for non-ASCII characters</p> <p>0.3.0 </p> <p>added year of publication; added separate files for languages: rus, pol, rum, cze, slo, gre with minor encoding issues</p> <p> </p> <p><em>Dataset Characteristics</em></p> <p>The following languages are available in separate data files:</p> <ul> <li>eng</li> <li>ger</li> <li>lat</li> <li>fre</li> <li>ita</li> <li>spa</li> <li>por</li> <li>dut</li> <li>swe</li> <li>dan</li> <li>nor</li> <li>ice</li> <li>fry</li> <li>rus*</li> <li>pol*</li> <li>rum*</li> <li>cze*</li> <li>slo*</li> <li>gre*</li> </ul> <p>*: Language file might be subject to character encoding issues.</p> <p>The other languages are present in the data set but have not been separated, i.e., they are combined in one data file:</p> <p>'fre', 'rus', 'pol', 'ger', 'eng', 'lit', 'dan', 'dut', 'spa', 'swe', 'ita', 'lat', 'nor', 'ind', 'bul', 'grc', 'fry', 'rum', 'cze', 'slo', 'bel', 'ice', 'fin', 'gre', 'hun', 'tur', 'enm', 'hrv', 'est', 'srp', 'roh', 'syr', 'wen', 'mal', 'afr', 'slv', 'mac', 'smi', 'nds', 'qmw', 'pra', 'oci', 'bre', 'san', 'alb', 'baq', 'non', 'ara', 'chm', 'per', 'cat', 'gmh', 'sla', 'arm', 'ukr', 'por', 'chu', 'heb', 'arc', 'gle', 'tib', 'lav', 'geo', 'crp', 'hin', 'mul', 'chi', 'epo', 'kor', 'kan', 'vot', 'csb', 'glg', 'kaz', 'frm', 'jpn', 'bur', 'srd', 'sal', 'ira', 'bos', 'mol', 'rom', 'tat', 'aze', 'yid', 'mar', 'mak', 'pli', 'rys', 'tgk', 'map', 'vie', 'tuk', 'oss', 'ota', 'tut', 'ben', 'sun', 'tir', 'bak', 'chv', 'ber', 'khm', 'may', 'pan', 'uzb', 'swa', 'kir', 'egy', 'dum', 'nep', 'cop', 'mon', 'tam', 'urd', 'zxx', 'wel', 'mis', 'ng', 'goh', 'dt', 'en', 'fao', 'fro', 'pus', 'kur', 'cus', 'hau', 'uig', 'sit', 'dt.', 'cpf', 'tgl', 'qoj', 'tag', 'raj', 'fiu', 'xal', 'kbd', 'udm', 'scr', 'gag', 'kas', 'scc', 'pro', 'tha', 'dar', 'dr', 'sna', 'ewe', 'de', 'dra', 'ang', 'ine', 'zza', 'und', 'ave', 'amh', 'crh', 'jav', 'cpe', 'akk', 'dsb', 'qce', 'guj', 'ltz', 'got', 'bua', 'peo', 'mdr', 'nob', 'ava', 'che', 'sux', 'kok', 'zap', 'nl', 'inc', 'sah', 'gem', 'law', 'bem', 'sin', 'qdo', 'hsb', 'som', 'lao', 'kam', 'kom', 'abk', 'roa', 'cau', 'ady', 'bat', 'mlt', 'sai', 'xho', 'paa', 'sot', 'bnt', 'lug', 'myn', 'kar', 'qhe', 'kin', 'zul', 'tsn', 'apa', 'nso', 'yao', 'yor', 'bih', 'nog', 'nap', 'loz', 'nbl', 'kon', 'nya', 'snh', 'chn', 'run', 'suk', 'fur', 'osa', 'bra', 'den', 'kpe', 'kal', 'tig', 'wol', 'gla', 'lad', 'mos', 'cre', 'krc', 'ge', 'fr', 'dak', 'fij', 'mad', 'srr', 'kum', 'her', 'nai', 'cel', 'inh', 'kro', 'hit', 'pal', 'tmh', 'tsw', 'bam', 'kab', 'kik', 'kua', 'lub', 'luo', 'nub', 'tem', 'znd', 'mai', 'tai', 'qkr', 'ful', 'man', 'lol', 'sag', 'tog', 'hai', 'arg', 'fat', 'nav', 'niu', 'ibo', 'ido', 'men', 'qju', 'gaa', 'vol', 'nah', 'mlg', 'nic', 'ijo', 'sus', 'orm', 'smo', 'mag', 'tyv', 'mnc', 'cos', 'mdf', 'kaa', 'dua', 'gez', 'ton', 'ven', 'snd', 'syc', 'nym', 'nia', 'sem', 'chg', 'fan', 'twi', 'mas', 'ina', 'ile', 'art', 'ori', 'qai', 'arw', 'mao', 'bas', 'kmb', 'tiv', 'bal', 'tar', 'tpi', 'abs', 'asm', 'qqa', 'iku', 'min', 'rup', 'tel', 'or', 'tah', 'aka', 'day', 'qqg', 'lah', 'lus', 'sio', 'oto', 'alg', 'shn', 'ndo', 'haw', 'tso', 'mus', 'cai', 'qev', 'new', 'zha', 'grn', 'khi', 'ssw', 'nde', 'bla', 'grb', 'mun', 'din', 'sam', 'mwr', 'cor', 'sat', 'cho', 'ger,', 'que', 'btk', 'glv', 'rar', 'jk', 'nno', 'cmc', 'mga', 'jw', 'iro', 'sog', 'hat', 'dzo', 'mkh', 'bik', 'ban', 'ilo', 'pam', 'ts', 'sme', 'myv', 'qnn', 'jpr', 'qte', 'yap', 'bis', 'sga', 'qkj', 'pap', 'ath', 'ipk', 'phi', 'sco', 'del', 'moh', 'iri', 'gae', 'ryl', 'our', 't--', 'grk', 'ssa', 'awa', 'efi', 'jrb', 'enk', 'kru', 'oji', 'arn', 'car', 'gsw', 'lez', 'war', 'ace', 'qrn', 'wln', 'ceb', 'aar', 'bug', 'kaw', 'chr', 'cpp', 'tet', 'aym', 'ces', 'hmo'</p>
Title, Author, Publisher, Place of Publication, and Language-related Network Graphs of the Berlin State Library Main Catalog
<p>The dataset contains graphs in GML, GraphML, and a simple JSON format.</p> <p>For each of the following languages:</p> <ol> <li>cze</li> <li>dan</li> <li>dut</li> <li>eng</li> <li>fre</li> <li>fry</li> <li>ger</li> <li>gre</li> <li>ice</li> <li>ita</li> <li>lat</li> <li>nor</li> <li>pol</li> <li>por</li> <li>rum</li> <li>rus</li> <li>slo</li> <li>spa</li> <li>swe</li> </ol> <p>two graphs are made available linking</p> <ul> <li>author, publisher, and place of publication</li> <li>author, publisher, place of publication, and title</li> </ul> <p>Additionaly, a third graph links authors and publishers to the language of publication (incl. year of the publication).</p> <p>The core statistics of each graph are outlined in <em>social_analysis_statistics.csv</em>. The smallest graph (fry, author_publisher_location) has 298 nodes and 264 edges, while the largest (ger, author_publisher_location_title) has 2,499,943 nodes and 3,950,900 edges.</p> <p>The language graphs spans all languages and has 1,706,273 nodes and 1,827,759 edges.</p> <p>All graphs have been created by a Python script available <a href="https://github.com/elektrobohemian/CulturalAnalytics/blob/master/SocialAnalysisStabikat.ipynb">here.</a></p>
Extracted Illustrations of the Berlin State Library's Digitized Collections (part 1 of 4)
<p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library's Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><em>Change Log for Version</em></p> <ol> <li>original dataset</li> <li>added color histograms (RGB, separated by channel) in JSON and Python pickle format as extracted by the <a href="https://pillow.readthedocs.io/en/stable/">Pillow</a> package (see https://github.com/elektrobohemian/StabiHacks/tree/master/image-tools)</li> </ol> <p><strong>This is part 1 of 4. The following datasets contain the other ZIP files (8 files in total):</strong></p> <ul> <li><a href="https://doi.org/10.5281/zenodo.2598145">https://doi.org/10.5281/zenodo.2598145</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598261">https://doi.org/10.5281/zenodo.2598261</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598270">https://doi.org/10.5281/zenodo.2598270</a></li> </ul> <p> </p>
OCR fulltexts of the Digital Collections of the Berlin State Library (DC-SBB)
<p>The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.</p> <p>At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.<br> For each page with OCR text, the language has been determined by <em>langid </em>(Lui/Baldwin 2012).</p> <p>corpus-entropy.pkl entropy rate per document page</p> <p>corpus-language.pkl language per document page</p> <p>corpus.zip fulltext corpus (extracts to .txt format)</p> <p>de_corpus.zip German sub-corpus (extracts to .txt format)</p> <p>selection_de.pkl Selection list of German documents</p> <p>xml2csv_alto.csv fulltext corpus per document page (incl.OCR word confidences)</p> <p> </p> <p><em>Sources</em></p> <p>Marco Lui and Timothy Baldwin. 2012. Langid.py:</p> <p>An off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations,</p> <p>ACL ’12, pages 25–30, Stroudsburg, PA, USA. Association for Computational Linguistics</p>
A reference library for Canadian invertebrates with 1.5 million barcodes, voucher specimens, and DNA samples
<p>[This repository contains the source data for the manuscript "A reference library for Canadian invertebrates with 1.5 million barcodes, voucher specimens, and DNA samples" by deWaard et al., 2019. BioRxiv]</p>
Machine-Readable Vocabulary Files of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)
<p>This dataset contains two versions of vocabulary files of the <a href="https://ark.staatsbibliothek-berlin.de/">ARK (Alter Realkatalog)</a> in .tsv and .ttl format used for training models for automatic subject indexing with the modular <a href="https://github.com/NatLibFi/Annif">Annif</a> tool. As the ARK is a historical classification system which has been used to describe historical works in the Staatsbibliothek zu Berlin – Berlin State Library’s collections up to 1955, this dataset has been created for generating automatic indexing suggestions for historical texts which have not yet been manually classified with the help of the ARK (for a detailed description of the ARK, see also <a href="../doi/10.5281/zenodo.12783813">Metadata of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)</a>. Together with specific corpus training data, these vocabulary files serve as input to Annif, with which the corresponding models on <a href="https://huggingface.co/SBB">Hugging Face at the Staatsbibliothek zu Berlin – Preußischer Kulturbesitz</a> community have been created. Associated corpus training data have been extracted from the <a href="../doi/10.5281/zenodo.12783813" target="_blank" rel="noopener">Metadata of the "Alter Realkatalog" (ARK)</a> (title data).</p>
EU volume, increment, aboveground biomass and BCEF libraries
<div> <p>This collection includes, for each EU-27 Member State (excluded Malta and Cyprus) the following database:</p> <div><strong>1. Volume and increment database</strong> (<em>Volume_increment_databse</em>) reporting the merchantable standing volume (in m<sup>3</sup> ha<sup>-1 </sup>under bark) and the net merchantable volume increment (in m<sup>3</sup> ha<sup>-1 </sup>yr<sup>-1 </sup>under bark) of the main forest types identified at country level.</div> <div><strong>2.</strong> <strong>Aboveground biomass and BCEF data</strong> (<em>Volume_biomass_bcef_databse</em>) reporting the total aboveground dry biomass (in t ha<sup>-1</sup>) of the main forest types identified at country level. The total aboveground dry biomass is further distinguished between stem (i.e., merchantable wood biomass excluding bark), bark, branches (including both main and small branches) and foliages. The database also reports the Biomass Conversion and Expansion factors (BCEF, adimensional) derived as the ratio between the merchantable volume associated to each record and the corresponding total aboveground biomass.</div> <div><strong>3.</strong> <strong> Ancillary data</strong> (<em>Volume_biomass_selection</em>) reporting detailed information on the parameters defining the allometric equations applicable to each forest type and country specific combinations of classifiers. A second file (<em>vol_to_biomass_best_match</em>) reports the RMSE values and the proportion of stemwood to total aboveground biomass associated to the allometric equation associated to each forest type and group of classifiers.</div> <p>All previous data were further scaled, when possible, at sub-national level, and distinguished between forest types (FT, defined according to the leading species reported by National Forest Inventories), management types (MT, mostly distinguishing high-forests and coppices) and management strategies (MS, eventually distinguishing even-aged and uneven-aged forest stands). All data, derived from countries' National Forest Inventory or other public datasources available at country level, were preliminarily harmonized according to a common definition of volume and increment.</p> </div>
vPro-MS peptide spectral library for the identification of human-pathogenic viruses by untargeted proteomics
<p>The viral proteomics workflow (vPro-MS) enables identification of human-pathogenic viruses from patient samples by untargeted proteomics. vPro-MS is based on an in-silico derived peptide library covering the human virome in <a href="https://www.uniprot.org/" rel="nofollow">UniProtKB</a> (331 viruses, 20,386 genomes, 121,977 peptides). vPro-MS is intended to identify human-pathogenic viruses from DiaNN (<a href="https://github.com/vdemichev/DiaNN">https://github.com/vdemichev/DiaNN</a>) outputs of either DIA or diaPASEF data. A scoring algorithm (vProID) assesses the confidence of virus identification and the results are finally summarized in a report table. </p> <p>The vPro Peptide Library folder contains 3 peptide FASTA files (Contaminants.fasta, Human.fasta, vPro.Virus.fasta), which were used to predict the spectral library (vPro-lib.predicted.speclib). Please note, that the additional commands “--cut” and “--duplicate-proteins” are needed to reprocess the prediction in DiaNN. This spectral library should be used to identify peptide sequences from samples of human origin using DiaNN. Furthermore, the folder contains the metadata file of the viral peptide sequences (vPro.Peptide.Library.txt) and a summary file of the virus taxonomy covered by the library (Taxonomy.Summary.txt). The metadata file is used by the vPro script to identify viruses from the DiaNN main report.</p>
Verification of library complexity in the HEK-Cas9 sublibraries - sequence data of the generated sublibraries A and B
<p>Sequence data of the generated HEK-Cas9 sublibraries A and B, linked to the manuscript 10.1128/mbio.01925-24: The <em>Bordetella</em> effector protein BteA induces host cell death by disruption of calcium homeostasis by Martin Zmuda, Eliska Sedlackova, Barbora Pravdova, Monika Cizkova, Marketa Dalecka, Ondrej Cerny, Tania Romero Allsop, Tomas Grousl, Ivana Malcova, and Jana Kamanova</p>
British Library Books genre detection model
<p><strong>Model description</strong></p> <p>This model is intended to predict, from the title of a book, whether it is 'fiction' or 'non-fiction'.</p> <p>This model was trained on data created from the <a href="https://www.bl.uk/collection-guides/digitised-printed-books">Digitised printed books (18th-19th Century)</a> book collection. The datasets in this collection are comprised and derived from 49,455 digitised books (65,227 volumes), mainly from the 19th Century. This dataset is dominated by English language books and includes books in several other languages in much smaller numbers. </p> <p>This model was originally developed for use as part of the <a href="https://livingwithmachines.ac.uk/">Living with Machines</a> project to be able to 'segment' this large dataset of books into different categories based on a 'crude' classification of genre i.e. whether the title was `fiction` or `non-fiction`.</p> <p>The model's training data (discussed more below) primarily consists of 19th Century book titles from the <a href="https://www.bl.uk/collection-guides/digitised-printed-books">British Library Digitised printed books (18th-19th century)</a> collection. These books have been catalogued according to British Library cataloguing practices. The model is likely to perform worse on any book titles from earlier or later periods. While the model is multilingual, it has training data in non-English book titles; these appear much less frequently.</p> <p><strong>How to use</strong></p> <p>To use this within fastai, first install version 2 of the fastai library. Following the documentation <a href="https://docs.fast.ai/#Installing">instructions</a>. Once you have fastai installed, you can use the model as follows:</p> <pre><code class="language-python">from fastai.text.all import load_learner learn = load_learner("20210928-model.pkl") learn.predict("Oliver Twist")</code></pre> <p><strong>Limitations and bias</strong></p> <p>The model was developed based on data from the <a href="https://www.bl.uk/collection-guides/digitised-printed-books">British Library's Digitised printed books (18th-19th Century)</a> collection. This dataset is not representative of books from the period covered with biases towards certain types (travel) and a likely absence of books that were difficult to digitise.</p> <p>The formatting of the British Library books corpus titles may differ from other collections, resulting in worse performance on other collections. It is recommended to evaluate the performance of the model before applying it to your own data. Likely, this model won't perform well for contemporary book titles without further fine-tuning.</p> <p><strong>Training data</strong></p> <p>The training data for this model will be available from the British Libary Research Repository shortly.</p> <p>The training data was created using the Zooniverse platform. British Library cataloguers carried out the majority of the annotations used as training data. More information on the process of creating the training data will be available soon. </p> <p><strong>Training procedure</strong></p> <p>Model training was carried out using the fastai library version 2.5.2. </p> <p>The notebook using for training the model will be available at: https://github.com/Living-with-machines/bl-books-genre-prediction</p> <p><strong>Eval result</strong></p> <p>The model was evaluated on a held out test set:</p> <pre> precision recall f1-score support Fiction 0.91 0.88 0.90 296 Non-fiction 0.94 0.95 0.95 554 accuracy 0.93 850 macro avg 0.93 0.92 0.92 850 weighted avg 0.93 0.93 0.93 850</pre>
Development of a spectral library for the discovery of altered genomic events in Mycobacterium avium associated with virulence using mass spectrometry-based proteogenomic analysis
<p><em>Mycobacterium avium</em> is one of the prominent disease-causing bacteria in humans. It causes lymphadenitis, chronic and extrapulmonary, and disseminated infections in adults, children, and immunocompromised patients. <em>M. avium</em> has ~4,500 predicted protein-coding regions on an average, which can be helpful in discovering several variants at the proteome level. Many of them are potentially associated with virulence, thus identifying such proteins can be a helpful feature in the development of panel-based theranostics. In line with such a long-term goal, we carried out an in-depth proteomic analysis of <em>M. avium</em> with both data-dependent and data-independent acquisition methods. Further, a set of proteogenomic investigations were carried out using the protein database for <em>Mycobacterium tuberculosis,</em> and a genome six-frame translated database and a variant protein database of <em>M. avium</em>. A search of mass spectrometry data analysis against <em>M. avium</em> protein database resulted in the identification of 2,954 proteins. Further, proteogenomic analyses aided in the identification of 1,301 novel peptide sequences and correction of translation start sites for 15 proteins. At the end, we created a spectral library of <em>M. avium</em> proteins including novel genome search-specific peptides and variant peptides detected in this study. We validated the spectral library by a data-independent acquisition of the <em>M. avium</em> proteome. Thus, we present a <em>M. avium </em>spectral library of 29,033 peptide precursors supported by 0.4 million fragment ions for further use by the biomedical community.</p>
ACM-Digital Library (ACM)
<p>ACM-Digital Library (ACM): a subset of the ACM Digital Library with 24, 897 documents containing articles related to Computer Science. We considered only the first level of the taxonomy adopted by ACM, where each document is assigned to one of 11 classes.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_<k>.pkl: pandas DataFrame with k-cross validation partition.</p> <p>The .zip contains all aforementioned files + the tfidf representation in the CSR matrix format.</p>
CoCO2 WP4.1 Library of Plumes
<p>This data collects synthetic CO2M images (the 'Library of Plumes' for CoCO2 WP4.1) in the form of NetCDF files, for 5 modelling systems applied to 7 case studies. The files contain 'raw' results (e.g., "CO2_PP_M") in units of mol/m2, total column results (e.g., "XCO2_PP_M") in units of ppm, and additionally the column mass ("dry_mol_mass" and surface pressure ("surface pressure").</p> <p>This work was carried out in the context of the EU project CoCO2, by the following institutes: Empa (Switzerland), Deutsche Wetterdienst (Germany), Le Laboratoire des Sciences du Climat et de l'Environnement (France), TNO (The Netherlands), Wageningen University & Research (the Netherlands). More details about the model runs, their input data, possible issues, expected data quality, etc., can be found in the accompanying report CoCO2 D4.2, <a href="https://www.coco2-project.eu/node/357">https://www.coco2-project.eu/node/357</a>.</p>
Diachronic and diatopic word embeddings from newspapers digitised by the British Library (1830-1889): North and South England
<p>Diachronic word embeddings (decade-level) trained with Word2Vec (via Gensim) on different geographic subcorpora of the Heritage Made Digital British and the Living with Machines historical newspaper collections:</p> <p>- North England (north.zip)</p> <p>- South England (south.zip)</p> <p>At the moment, for each subcorpus, Word2Vec models are available for each decade in the period 1830-1889. More models are on the way for the following:</p> <p>- each decade in the periods 1780-1829 and 1890-1920 for both North and South England.</p> <p>- diachronic models for the following regions: Scotland, Wales, and Midlands.</p> <p>The models were trained using the following parameters:</p> <pre><code>sg = True min_count = 1 window = 5 vector_size = 200 epochs = 5</code></pre> <p>Like the embeddings in <a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, the model for each decade was aligned to the most recent one with Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation: <a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a>.</p> <p>Project website (Living with Machines): <a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p> <p>Data related to: Nilo Pedrazzini & Barbara McGillivray, <em>Diachronic and diatopic word embeddings from British historical newspapers</em>, presented at AIUCD (Convegno dell’Associazione per l’Informatica Umanistica e la Cultura Digitale) in Siena (Italy), June 2023.</p>
Spectral Library of European Pegmatites, Pegmatite Minerals and Pegmatite Host-Rocks – The Greenpeg Database
<p>Spectral signature, obtained through reflectance spectroscopy studies, of European pegmatites and minerals, as well of their host rocks. Samples include LCT- and NYF-type pegmatites and host rocks from pegmatite locations in Austria, Ireland, Norway, Portugal, and Spain. Sample preparation and spectral measurement were conducted in the Universidade do Porto – Faculdade de Ciências (UPORTO) laboratories. The database contains the reflectance spectra (raw and with continuum removed), sample photographs, and main absorption features automatically extracted by a Python routine. Whenever possible, spectral mineralogy was interpreted based on the continuum-removed spectra. A detailed description of the database, its content, the measuring instrument, and interoperability with GIS is found in the database report.</p>
[ARAMCO Dhahran library collection 14430 AMK] ٱلْمَمْلَكَة ٱلْعَرَبِيَّة ٱلسُّعُوْدِيَّة. ٱلْهُفُوف
<p>[ARAMCO 14430 AMK] Kingdom of Saudi Arabia, al-Hufūf. Oblique aerial photograph of town and al-Kut, before 1973.</p>
[ARAMCO Dhahran library collection 9051 TFW] ٱلْمَمْلَكَة ٱلْعَرَبِيَّة ٱلسُّعُوْدِيَّة. ٱلْهُفُوف .الكوت. قصرالأمير
<p>[ARAMCO 9051 TFW] Kingdom of Saudi Arabia, al-Hufūf. Emir's palace, as documented in the 1960s, now demolished. The current أمارة محافظة الأحساء is standing on the site. Coordinates: 25°22'39"N 49°35'13"E.</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.