Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,307

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,307 results for “libraries”

Learn how ShareScore rates datasets ↗
zenodo44/100

Developer Expertise Dataset on JavaScript Libraries

<p>This dataset contains an anonymized list of surveyed developers who&nbsp;provided their expertise level on three popular JavaScript libraries:</p> <ol> <li><a href="https://github.com/facebook/react">ReactJS</a>, a library for building enriched web interfaces&nbsp;</li> <li><a href="https://github.com/mongodb/node-mongodb-native">MongoDB</a>, a driver for accessing MongoDB databased&nbsp;</li> <li><a href="https://github.com/socketio/socket.io">Socket.IO</a>, a library for realtime communication&nbsp;&nbsp;</li> </ol>

opencc-by-4.0Nov 2018View details →
zenodo44/100

GALAKSIENN: the G.A.S semi-analytical model associtated library

<p>The GALAKSIENN library contains data sets generated by the G.A.S. semi-analytical model of galaxy formation and evolution. The G.A.S. model is fully described in a set of three Astronomy &amp; Astrophysics papers: G.A.S. I:&nbsp; A prescription for turbulence-regulated star formation and its impact on galaxy properties; G.A.S. II: Dust extinction in galaxies, luminosity functions and InfraRed Excess and G.A.S. III: The panchromatic view of galaxies, Stellar/dust continua and main gas emission lines. The library stores mock galaxy catalogues and ascii tables (stellar mass functions, luminosity functions, number counts ...).</p>

opencc-by-4.0Nov 2018View details →
zenodo44/100

Metadata, Title Pages, and Network Graph of the Digitized Content of the Berlin State Library (146,000 items)

<p>The data set has been downloaded via the OAI-PMH endpoint of the Berlin State Library/Staatsbibliothek zu Berlin&rsquo;s Digitized Collections (<a href="https://digital.staatsbibliothek-berlin.de/oai">https://digital.staatsbibliothek-berlin.de/oai</a>) on March 1<sup>st</sup> 2019 and converted into common tabular formats on the basis of the provided Dublin Core metadata. It contains 146,000 records.</p> <p>In addition to the bibliographic metadata, representative images of the works have been downloaded, resized to a 512 pixel maximum thumbnail image and saved in JPEG format. The image data is split into title pages and first pages. Title pages have been derived from structural metadata created by scan operators and librarians. If this information was not available, first pages of the media have been downloaded. In case of multi-volume media, title pages are not available.</p> <p>In total, 141,206 images title/first pages are available.</p> <p>&nbsp;</p> <p>Furthermore, the tabular data has been cleaned and extended with geo-spatial coordinates provided by the OpenStreetMap project (<a href="https://www.openstreetmap.org">https://www.openstreetmap.org</a>). The actual data processing steps are summarized in the next section. For the sake of transparency and reproducibility, the original data taken from the OAI-PMH endpoint is still present in the table.</p> <p>&nbsp;</p> <p>To conclude with, various graphs in GML file format are available that can be loaded directly into graph analysis tools such as Gephi (<a href="https://gephi.org/">https://gephi.org/</a>).</p> <p>&nbsp;</p> <p>The implementation of the data processing steps (incl. graph creation) are available as a Jupyter notebook provided at <a href="https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb">https://github.com/elektrobohemian/SBBrowse2018/blob/master/DataProcessing.ipynb</a>.</p> <p>&nbsp;</p> <p>Tabular Metadata</p> <p>&nbsp;</p> <p>The metadata is available in Excel (cleanedData.xlsx) and CSV (cleanedData.csv) file formats with equal content.</p> <p>The table contains the following columns. Italique columns have not been processed.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>title</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The title of the medium</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>creator</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Its creator (family name, first name)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>subject</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A collection&rsquo;s name as provided by the library</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>type</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The type of medium</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>format</em> &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A MIME type for full metadata download</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>identifier</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; An additional identifier (most often the PPN)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>language</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A 3-letter language code of the medium</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>date</em>&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The date of creation/publication or a time span</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>relation</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A relation to a project or collection a medium has been digitized for.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>coverage</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The location of publication or origin (ranging from cities to continents)</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>publisher</em>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The publisher of the medium.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>rights</em>&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Copyright information.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; <em>PPN</em>&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The unique identifier that can be used to find more information about the current medium in all information systems of Berlin State Library/Staatsbibliothek zu Berlin.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; spatialClean&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; In case of multiple entries in coverage, only the first place of origin has been extracted. Additionally, characters such as question marks, brackets, or the like have been removed. The entries have been normalized regarding whitespaces and writing variants with the help of regular expressions.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; dateClean&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; As the original date may contain various format variants to indicate unclear creation dates (e.g., time spans or question marks), this field contains a mapping to a certain point in time.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; spatialCluster &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The cluster ID determined with the help of the Jaro-Winkler distance on the spatialClean string. This step is needed because the spatialClean fields still contain a huge amount of orthographic variants and latinizations of geographic names.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; spatialClusterName&nbsp;&nbsp; A verbal cluster name (controlled manually).</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; latitude&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The latitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; longitude&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The longitude provided by OpenStreetMap of the spatialClusterName if the location could be found.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; century&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A century derived from the date.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; textCluster&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A text cluster ID on the basis of a k-means clustering relying on the title field with a vocabulary size of 125,000 using the tf*idf model and k=5,000.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; creatorCluster &nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; A text cluster ID based on the creator field with k=20,000.</p> <p>&middot;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; titleImage&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; The path to the first/title page relative to the img/ subdirectory or None in case of a multi-volume work.</p> <p>Other Data</p> <p>&nbsp;</p> <p><em>graphs.zip</em></p> <p>&nbsp;</p> <p>Various pre-computed graphs.</p> <p><em>&nbsp;</em></p> <p><em>img.zip</em></p> <p>&nbsp;</p> <p>First and title pages in JPEG format.</p> <p>&nbsp;</p> <p><em>json.zip</em></p> <p>&nbsp;</p> <p>JSON files for each record in the following format:</p> <p>&nbsp;</p> <p>ppn&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;PPN57346250X&quot;</p> <p>dateClean&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;1625&quot;</p> <p>title&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;M. Georgii Gutkii, Gymnasii Berlinensis Rectoris Habitus Primorum Principiorum, Seu Intelligentia; Annexae Sunt Appendicis loco Disputationes super eodem habitu tum in Academia Wittebergensi, tum in Gymnasio Berlinensi ventilatae&quot;</p> <p>creator&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Gutke, Georg&quot;</p> <p>spatialClusterName&nbsp;&nbsp; &quot;Berlin&quot;</p> <p>spatialClean&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Berolini&quot;</p> <p>spatialRaw&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Berolini&quot;</p> <p>mediatype&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;monograph&quot;</p> <p>subject&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Historische Drucke&quot;</p> <p>publisher&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;Kallius&quot;</p> <p>lat&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;52.5170365&quot;</p> <p>lng&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;13.3888599&quot;</p> <p>textCluster&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;45&quot;</p> <p>creatorCluster&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;5040&quot;</p> <p>titleImage&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &quot;titlepages/PPN57346250X.jpg&quot;</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Extract from the Berlin State Library's Main Catalog

<p>The data set is based on the main catalog of the library. Currently, the following fields are extracted:</p> <ul> <li>title</li> <li>author (+ optional GND ID)</li> <li>publisher</li> <li>place of publication</li> <li>country of publication</li> <li>year of publication</li> </ul> <p>The extract has been created by the processPicaPlus script available <a href="https://github.com/elektrobohemian/StabiHacks">here</a>. Attention, some special characters might not have been extracted correctly in versions &lt;1.0.0.</p> <p><em>Change Log:</em></p> <p>0.2.0&nbsp;&nbsp;&nbsp;</p> <p>fixes various encoding issues for non-ASCII characters</p> <p>0.3.0&nbsp;&nbsp;&nbsp;</p> <p>added year of publication; added separate files for languages: rus, pol, rum, cze, slo, gre with minor encoding issues</p> <p>&nbsp;</p> <p><em>Dataset Characteristics</em></p> <p>The following languages are available in separate data files:</p> <ul> <li>eng</li> <li>ger</li> <li>lat</li> <li>fre</li> <li>ita</li> <li>spa</li> <li>por</li> <li>dut</li> <li>swe</li> <li>dan</li> <li>nor</li> <li>ice</li> <li>fry</li> <li>rus*</li> <li>pol*</li> <li>rum*</li> <li>cze*</li> <li>slo*</li> <li>gre*</li> </ul> <p>*: Language file might be subject to character encoding issues.</p> <p>The other languages are present in the data set but have not been separated, i.e., they are combined in one data file:</p> <p>&#39;fre&#39;, &#39;rus&#39;, &#39;pol&#39;, &#39;ger&#39;, &#39;eng&#39;, &#39;lit&#39;, &#39;dan&#39;, &#39;dut&#39;, &#39;spa&#39;, &#39;swe&#39;, &#39;ita&#39;, &#39;lat&#39;, &#39;nor&#39;, &#39;ind&#39;, &#39;bul&#39;, &#39;grc&#39;, &#39;fry&#39;, &#39;rum&#39;, &#39;cze&#39;, &#39;slo&#39;, &#39;bel&#39;, &#39;ice&#39;, &#39;fin&#39;, &#39;gre&#39;, &#39;hun&#39;, &#39;tur&#39;, &#39;enm&#39;, &#39;hrv&#39;, &#39;est&#39;, &#39;srp&#39;, &#39;roh&#39;, &#39;syr&#39;, &#39;wen&#39;, &#39;mal&#39;, &#39;afr&#39;, &#39;slv&#39;, &#39;mac&#39;, &#39;smi&#39;, &#39;nds&#39;, &#39;qmw&#39;, &#39;pra&#39;, &#39;oci&#39;, &#39;bre&#39;, &#39;san&#39;, &#39;alb&#39;, &#39;baq&#39;, &#39;non&#39;, &#39;ara&#39;, &#39;chm&#39;, &#39;per&#39;, &#39;cat&#39;, &#39;gmh&#39;, &#39;sla&#39;, &#39;arm&#39;, &#39;ukr&#39;, &#39;por&#39;, &#39;chu&#39;, &#39;heb&#39;, &#39;arc&#39;, &#39;gle&#39;, &#39;tib&#39;, &#39;lav&#39;, &#39;geo&#39;, &#39;crp&#39;, &#39;hin&#39;, &#39;mul&#39;, &#39;chi&#39;, &#39;epo&#39;, &#39;kor&#39;, &#39;kan&#39;, &#39;vot&#39;, &#39;csb&#39;, &#39;glg&#39;, &#39;kaz&#39;, &#39;frm&#39;, &#39;jpn&#39;, &#39;bur&#39;, &#39;srd&#39;, &#39;sal&#39;, &#39;ira&#39;, &#39;bos&#39;, &#39;mol&#39;, &#39;rom&#39;, &#39;tat&#39;, &#39;aze&#39;, &#39;yid&#39;, &#39;mar&#39;, &#39;mak&#39;, &#39;pli&#39;, &#39;rys&#39;, &#39;tgk&#39;, &#39;map&#39;, &#39;vie&#39;, &#39;tuk&#39;, &#39;oss&#39;, &#39;ota&#39;, &#39;tut&#39;, &#39;ben&#39;, &#39;sun&#39;, &#39;tir&#39;, &#39;bak&#39;, &#39;chv&#39;, &#39;ber&#39;, &#39;khm&#39;, &#39;may&#39;, &#39;pan&#39;, &#39;uzb&#39;, &#39;swa&#39;, &#39;kir&#39;, &#39;egy&#39;, &#39;dum&#39;, &#39;nep&#39;, &#39;cop&#39;, &#39;mon&#39;, &#39;tam&#39;, &#39;urd&#39;, &#39;zxx&#39;, &#39;wel&#39;, &#39;mis&#39;, &#39;ng&#39;, &#39;goh&#39;, &#39;dt&#39;, &#39;en&#39;, &#39;fao&#39;, &#39;fro&#39;, &#39;pus&#39;, &#39;kur&#39;, &#39;cus&#39;, &#39;hau&#39;, &#39;uig&#39;, &#39;sit&#39;, &#39;dt.&#39;, &#39;cpf&#39;, &#39;tgl&#39;, &#39;qoj&#39;, &#39;tag&#39;, &#39;raj&#39;, &#39;fiu&#39;, &#39;xal&#39;, &#39;kbd&#39;, &#39;udm&#39;, &#39;scr&#39;, &#39;gag&#39;, &#39;kas&#39;, &#39;scc&#39;, &#39;pro&#39;, &#39;tha&#39;, &#39;dar&#39;, &#39;dr&#39;, &#39;sna&#39;, &#39;ewe&#39;, &#39;de&#39;, &#39;dra&#39;, &#39;ang&#39;, &#39;ine&#39;, &#39;zza&#39;, &#39;und&#39;, &#39;ave&#39;, &#39;amh&#39;, &#39;crh&#39;, &#39;jav&#39;, &#39;cpe&#39;, &#39;akk&#39;, &#39;dsb&#39;, &#39;qce&#39;, &#39;guj&#39;, &#39;ltz&#39;, &#39;got&#39;, &#39;bua&#39;, &#39;peo&#39;, &#39;mdr&#39;, &#39;nob&#39;, &#39;ava&#39;, &#39;che&#39;, &#39;sux&#39;, &#39;kok&#39;, &#39;zap&#39;, &#39;nl&#39;, &#39;inc&#39;, &#39;sah&#39;, &#39;gem&#39;, &#39;law&#39;, &#39;bem&#39;, &#39;sin&#39;, &#39;qdo&#39;, &#39;hsb&#39;, &#39;som&#39;, &#39;lao&#39;, &#39;kam&#39;, &#39;kom&#39;, &#39;abk&#39;, &#39;roa&#39;, &#39;cau&#39;, &#39;ady&#39;, &#39;bat&#39;, &#39;mlt&#39;, &#39;sai&#39;, &#39;xho&#39;, &#39;paa&#39;, &#39;sot&#39;, &#39;bnt&#39;, &#39;lug&#39;, &#39;myn&#39;, &#39;kar&#39;, &#39;qhe&#39;, &#39;kin&#39;, &#39;zul&#39;, &#39;tsn&#39;, &#39;apa&#39;, &#39;nso&#39;, &#39;yao&#39;, &#39;yor&#39;, &#39;bih&#39;, &#39;nog&#39;, &#39;nap&#39;, &#39;loz&#39;, &#39;nbl&#39;, &#39;kon&#39;, &#39;nya&#39;, &#39;snh&#39;, &#39;chn&#39;, &#39;run&#39;, &#39;suk&#39;, &#39;fur&#39;, &#39;osa&#39;, &#39;bra&#39;, &#39;den&#39;, &#39;kpe&#39;, &#39;kal&#39;, &#39;tig&#39;, &#39;wol&#39;, &#39;gla&#39;, &#39;lad&#39;, &#39;mos&#39;, &#39;cre&#39;, &#39;krc&#39;, &#39;ge&#39;, &#39;fr&#39;, &#39;dak&#39;, &#39;fij&#39;, &#39;mad&#39;, &#39;srr&#39;, &#39;kum&#39;, &#39;her&#39;, &#39;nai&#39;, &#39;cel&#39;, &#39;inh&#39;, &#39;kro&#39;, &#39;hit&#39;, &#39;pal&#39;, &#39;tmh&#39;, &#39;tsw&#39;, &#39;bam&#39;, &#39;kab&#39;, &#39;kik&#39;, &#39;kua&#39;, &#39;lub&#39;, &#39;luo&#39;, &#39;nub&#39;, &#39;tem&#39;, &#39;znd&#39;, &#39;mai&#39;, &#39;tai&#39;, &#39;qkr&#39;, &#39;ful&#39;, &#39;man&#39;, &#39;lol&#39;, &#39;sag&#39;, &#39;tog&#39;, &#39;hai&#39;, &#39;arg&#39;, &#39;fat&#39;, &#39;nav&#39;, &#39;niu&#39;, &#39;ibo&#39;, &#39;ido&#39;, &#39;men&#39;, &#39;qju&#39;, &#39;gaa&#39;, &#39;vol&#39;, &#39;nah&#39;, &#39;mlg&#39;, &#39;nic&#39;, &#39;ijo&#39;, &#39;sus&#39;, &#39;orm&#39;, &#39;smo&#39;, &#39;mag&#39;, &#39;tyv&#39;, &#39;mnc&#39;, &#39;cos&#39;, &#39;mdf&#39;, &#39;kaa&#39;, &#39;dua&#39;, &#39;gez&#39;, &#39;ton&#39;, &#39;ven&#39;, &#39;snd&#39;, &#39;syc&#39;, &#39;nym&#39;, &#39;nia&#39;, &#39;sem&#39;, &#39;chg&#39;, &#39;fan&#39;, &#39;twi&#39;, &#39;mas&#39;, &#39;ina&#39;, &#39;ile&#39;, &#39;art&#39;, &#39;ori&#39;, &#39;qai&#39;, &#39;arw&#39;, &#39;mao&#39;, &#39;bas&#39;, &#39;kmb&#39;, &#39;tiv&#39;, &#39;bal&#39;, &#39;tar&#39;, &#39;tpi&#39;, &#39;abs&#39;, &#39;asm&#39;, &#39;qqa&#39;, &#39;iku&#39;, &#39;min&#39;, &#39;rup&#39;, &#39;tel&#39;, &#39;or&#39;, &#39;tah&#39;, &#39;aka&#39;, &#39;day&#39;, &#39;qqg&#39;, &#39;lah&#39;, &#39;lus&#39;, &#39;sio&#39;, &#39;oto&#39;, &#39;alg&#39;, &#39;shn&#39;, &#39;ndo&#39;, &#39;haw&#39;, &#39;tso&#39;, &#39;mus&#39;, &#39;cai&#39;, &#39;qev&#39;, &#39;new&#39;, &#39;zha&#39;, &#39;grn&#39;, &#39;khi&#39;, &#39;ssw&#39;, &#39;nde&#39;, &#39;bla&#39;, &#39;grb&#39;, &#39;mun&#39;, &#39;din&#39;, &#39;sam&#39;, &#39;mwr&#39;, &#39;cor&#39;, &#39;sat&#39;, &#39;cho&#39;, &#39;ger,&#39;, &#39;que&#39;, &#39;btk&#39;, &#39;glv&#39;, &#39;rar&#39;, &#39;jk&#39;, &#39;nno&#39;, &#39;cmc&#39;, &#39;mga&#39;, &#39;jw&#39;, &#39;iro&#39;, &#39;sog&#39;, &#39;hat&#39;, &#39;dzo&#39;, &#39;mkh&#39;, &#39;bik&#39;, &#39;ban&#39;, &#39;ilo&#39;, &#39;pam&#39;, &#39;ts&#39;, &#39;sme&#39;, &#39;myv&#39;, &#39;qnn&#39;, &#39;jpr&#39;, &#39;qte&#39;, &#39;yap&#39;, &#39;bis&#39;, &#39;sga&#39;, &#39;qkj&#39;, &#39;pap&#39;, &#39;ath&#39;, &#39;ipk&#39;, &#39;phi&#39;, &#39;sco&#39;, &#39;del&#39;, &#39;moh&#39;, &#39;iri&#39;, &#39;gae&#39;, &#39;ryl&#39;, &#39;our&#39;, &#39;t--&#39;, &#39;grk&#39;, &#39;ssa&#39;, &#39;awa&#39;, &#39;efi&#39;, &#39;jrb&#39;, &#39;enk&#39;, &#39;kru&#39;, &#39;oji&#39;, &#39;arn&#39;, &#39;car&#39;, &#39;gsw&#39;, &#39;lez&#39;, &#39;war&#39;, &#39;ace&#39;, &#39;qrn&#39;, &#39;wln&#39;, &#39;ceb&#39;, &#39;aar&#39;, &#39;bug&#39;, &#39;kaw&#39;, &#39;chr&#39;, &#39;cpp&#39;, &#39;tet&#39;, &#39;aym&#39;, &#39;ces&#39;, &#39;hmo&#39;</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Title, Author, Publisher, Place of Publication, and Language-related Network Graphs of the Berlin State Library Main Catalog

<p>The dataset contains graphs in GML, GraphML, and a simple JSON format.</p> <p>For each of the following languages:</p> <ol> <li>cze</li> <li>dan</li> <li>dut</li> <li>eng</li> <li>fre</li> <li>fry</li> <li>ger</li> <li>gre</li> <li>ice</li> <li>ita</li> <li>lat</li> <li>nor</li> <li>pol</li> <li>por</li> <li>rum</li> <li>rus</li> <li>slo</li> <li>spa</li> <li>swe</li> </ol> <p>two graphs are made available linking</p> <ul> <li>author, publisher, and place of publication</li> <li>author, publisher, place of publication, and title</li> </ul> <p>Additionaly, a third graph links authors and publishers to the language of publication (incl. year of the publication).</p> <p>The core statistics of each graph are outlined in <em>social_analysis_statistics.csv</em>. The smallest graph (fry, author_publisher_location) has 298 nodes and 264 edges, while the largest (ger, author_publisher_location_title) has 2,499,943 nodes and 3,950,900 edges.</p> <p>The language graphs spans all languages and has 1,706,273 nodes and 1,827,759 edges.</p> <p>All graphs have been created by a Python script available <a href="https://github.com/elektrobohemian/CulturalAnalytics/blob/master/SocialAnalysisStabikat.ipynb">here.</a></p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

Extracted Illustrations of the Berlin State Library's Digitized Collections (part 1 of 4)

<p>The dataset consists of various illustrations extracted from 26,233 historical books and other media offered in the Berlin State Library&#39;s Digitized Collections. The media objects are older than 1920.</p> <p>Version 1.0 contains of 594,890 extracted illustrations in total.</p> <p>The extraction of illustrations is driven by the coordinates given by the ABBYY FineReader OCR engine (in ALTO XML) . The extracted illustrations have not been resized but compressed and saved in JPEG format.</p> <p>Pre-trained models in order to separate color scales, hand-written signatures, library stamps or the like from interesting content are available under: <a href="https://github.com/elektrobohemian/imi-unicorns">https://github.com/elektrobohemian/imi-unicorns</a>.</p> <p>The extracts for each media object are stored in separated sub-folders and tar files named after the PPN (a unique ID used in the library) to facilitate further processing. Additional metadata can be obtained with help of the PPN as described here: <a href="https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md">https://github.com/elektrobohemian/StabiHacks/blob/master/ppn-howto.md</a> .</p> <p>The dataset is published as a set of ZIP files, each fitting on a Blu Ray disc. <strong>After decompression, the contents will consume ca. 166 GB.</strong></p> <p><em>Change Log for Version</em></p> <ol> <li>original dataset</li> <li>added color histograms (RGB, separated by channel) in JSON and Python pickle format as extracted by the <a href="https://pillow.readthedocs.io/en/stable/">Pillow</a> package (see https://github.com/elektrobohemian/StabiHacks/tree/master/image-tools)</li> </ol> <p><strong>This is part 1 of 4. The following datasets contain the other ZIP files (8 files in total):</strong></p> <ul> <li><a href="https://doi.org/10.5281/zenodo.2598145">https://doi.org/10.5281/zenodo.2598145</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598261">https://doi.org/10.5281/zenodo.2598261</a></li> <li><a href="https://doi.org/10.5281/zenodo.2598270">https://doi.org/10.5281/zenodo.2598270</a></li> </ul> <p>&nbsp;</p>

opencc-by-4.0Mar 2019View details →
zenodo44/100

OCR fulltexts of the Digital Collections of the Berlin State Library (DC-SBB)

<p>The digital collections of the SBB contain 153,942 digitized works from the time period of 1470 to 1945.</p> <p>At the time of publication, 28,909 works have been OCR-processed resulting in 4,988,099 full-text pages.<br> For each page with OCR text, the language has been determined by <em>langid </em>(Lui/Baldwin 2012).</p> <p>corpus-entropy.pkl &nbsp; &nbsp;&nbsp; entropy rate per document page</p> <p>corpus-language.pkl&nbsp;&nbsp; language per document page</p> <p>corpus.zip &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; fulltext corpus (extracts to .txt format)</p> <p>de_corpus.zip &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp; German sub-corpus (extracts to .txt format)</p> <p>selection_de.pkl&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; Selection list of German documents</p> <p>xml2csv_alto.csv&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; fulltext corpus per document page (incl.OCR word confidences)</p> <p>&nbsp;</p> <p><em>Sources</em></p> <p>Marco Lui and Timothy Baldwin. 2012. Langid.py:</p> <p>An off-the-shelf language identification tool. In Proceedings of the ACL 2012 System Demonstrations,</p> <p>ACL &rsquo;12, pages 25&ndash;30, Stroudsburg, PA, USA. Association for Computational Linguistics</p>

opencc-by-4.0Jun 2019View details →
zenodo44/100

A reference library for Canadian invertebrates with 1.5 million barcodes, voucher specimens, and DNA samples

<p>[This repository contains the source data for the manuscript &quot;A reference library for Canadian invertebrates with 1.5 million barcodes, voucher specimens, and DNA samples&quot; by deWaard et al., 2019. BioRxiv]</p>

opencc-zeroJun 2019View details →
zenodo44/100

Machine-Readable Vocabulary Files of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)

<p>This dataset contains two versions of vocabulary files of the&nbsp;<a href="https://ark.staatsbibliothek-berlin.de/">ARK (Alter Realkatalog)</a> in .tsv and .ttl format used for training models for automatic subject indexing with the modular <a href="https://github.com/NatLibFi/Annif">Annif</a> tool. As the ARK is a historical classification system which has been used to describe historical works in the Staatsbibliothek zu Berlin &ndash; Berlin State Library&rsquo;s collections up to 1955, this dataset has been created for generating automatic indexing suggestions for historical texts which have not yet been manually classified with the help of the ARK (for a detailed description of the ARK, see also <a href="../doi/10.5281/zenodo.12783813">Metadata of the "Alter Realkatalog" (ARK) of Berlin State Library (SBB)</a>. Together with specific corpus training data, these vocabulary files serve as input to Annif, with which the corresponding models on <a href="https://huggingface.co/SBB">Hugging Face at the Staatsbibliothek zu Berlin &ndash; Preu&szlig;ischer Kulturbesitz</a> community have been created. Associated corpus training data have been extracted from the <a href="../doi/10.5281/zenodo.12783813" target="_blank" rel="noopener">Metadata of the "Alter Realkatalog" (ARK)</a> (title data).</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

EU volume, increment, aboveground biomass and BCEF libraries

<div> <p>This collection includes, for each EU-27 Member State (excluded Malta and Cyprus) the following database:</p> <div><strong>1. Volume and increment database</strong> (<em>Volume_increment_databse</em>) reporting the merchantable standing volume (in m<sup>3</sup> ha<sup>-1 </sup>under bark) &nbsp;and the net merchantable volume increment (in m<sup>3</sup> ha<sup>-1 </sup>yr<sup>-1 </sup>under bark) &nbsp;of the main forest types identified at country level.</div> <div><strong>2.</strong> <strong>Aboveground biomass and BCEF data</strong> (<em>Volume_biomass_bcef_databse</em>) reporting the total aboveground dry biomass (in t ha<sup>-1</sup>) of the main forest types identified at country level. The total aboveground dry biomass is further distinguished between stem (i.e., merchantable wood biomass excluding bark), bark, branches (including both main and small branches) and foliages. &nbsp;The database also reports the Biomass Conversion and Expansion factors (BCEF, adimensional) derived as the ratio between the merchantable volume associated to each record and the corresponding total aboveground biomass.</div> <div><strong>3.</strong> <strong> Ancillary data</strong> (<em>Volume_biomass_selection</em>) reporting &nbsp;detailed information on the parameters defining the allometric equations applicable to each forest type and country specific combinations of classifiers. A second file (<em>vol_to_biomass_best_match</em>) reports the RMSE values and the proportion of stemwood to total aboveground biomass associated to the allometric equation associated to each forest type and group of classifiers.</div> <p>All previous data were further scaled, when possible, at sub-national level, and distinguished between forest types (FT, defined according to the leading species reported by National Forest Inventories), management types (MT, mostly distinguishing high-forests and coppices) and management strategies (MS, eventually distinguishing even-aged and uneven-aged forest stands). &nbsp;All data, derived from countries' National Forest Inventory or other public datasources available at country level, were preliminarily harmonized according to a common definition of volume and increment.</p> </div>

opencc-by-4.0May 2024View details →
zenodo44/100

vPro-MS peptide spectral library for the identification of human-pathogenic viruses by untargeted proteomics

<p>The viral proteomics workflow (vPro-MS) enables identification of human-pathogenic viruses from patient samples by untargeted proteomics. vPro-MS is based on an in-silico derived peptide library covering the human virome in <a href="https://www.uniprot.org/" rel="nofollow">UniProtKB</a> (331 viruses, 20,386 genomes, 121,977 peptides). vPro-MS is intended to identify human-pathogenic viruses from DiaNN (<a href="https://github.com/vdemichev/DiaNN">https://github.com/vdemichev/DiaNN</a>) outputs of either DIA or diaPASEF data. A scoring algorithm (vProID) assesses the confidence of virus identification and the results are finally summarized in a report table.&nbsp;</p> <p>The vPro Peptide Library folder contains 3 peptide FASTA files (Contaminants.fasta, Human.fasta, vPro.Virus.fasta), which were used to predict the spectral library (vPro-lib.predicted.speclib). Please note, that the additional commands &ldquo;--cut&rdquo; and &ldquo;--duplicate-proteins&rdquo; are needed to reprocess the prediction in DiaNN. This spectral library should be used to identify peptide sequences from samples of human origin using DiaNN. Furthermore, the folder contains the metadata file of the viral peptide sequences (vPro.Peptide.Library.txt) and a summary file of the virus taxonomy covered by the library (Taxonomy.Summary.txt). The metadata file is used by the vPro script to identify viruses from the DiaNN main report.</p>

opencc-by-4.0Sep 2024View details →
zenodo44/100

Verification of library complexity in the HEK-Cas9 sublibraries - sequence data of the generated sublibraries A and B

<p>Sequence data of the generated HEK-Cas9 sublibraries A and B, linked to the manuscript 10.1128/mbio.01925-24: The <em>Bordetella</em> effector protein BteA induces host cell death by disruption of calcium homeostasis by Martin Zmuda, Eliska Sedlackova, Barbora Pravdova, Monika Cizkova, Marketa Dalecka, Ondrej Cerny, Tania Romero Allsop, Tomas Grousl, Ivana Malcova, and Jana Kamanova</p>

opencc-by-4.0Nov 2024View details →
zenodo44/100

British Library Books genre detection model

<p><strong>Model description</strong></p> <p>This model is intended to predict, from the title of a book, whether it is &#39;fiction&#39; or &#39;non-fiction&#39;.</p> <p>This model was trained on data created from the <a href="https://www.bl.uk/collection-guides/digitised-printed-books">Digitised printed books (18th-19th Century)</a> book collection. The datasets in this collection are comprised and derived from 49,455 digitised books (65,227 volumes), mainly from the 19th Century. This dataset is dominated by English language books and includes books in several other languages in much smaller numbers.&nbsp;</p> <p>This model was originally developed for use as part of the <a href="https://livingwithmachines.ac.uk/">Living with Machines</a> project to be able to &#39;segment&#39; this large dataset of books into different categories based on a &#39;crude&#39; classification of genre i.e. whether the title was `fiction` or `non-fiction`.</p> <p>The model&#39;s training data (discussed more below) primarily consists of 19th Century book titles from the <a href="https://www.bl.uk/collection-guides/digitised-printed-books">British Library&nbsp;Digitised printed books (18th-19th century)</a> collection. These books have been catalogued according to British Library cataloguing practices. The model is likely to perform worse on any book titles from earlier or later periods. While the model is multilingual, it has training data in non-English book titles; these appear much less frequently.</p> <p><strong>How to use</strong></p> <p>To use this within fastai, first install version 2 of the fastai library. Following the documentation&nbsp;<a href="https://docs.fast.ai/#Installing">instructions</a>. Once you have fastai installed, you can use the model as follows:</p> <pre><code class="language-python">from fastai.text.all import load_learner learn = load_learner("20210928-model.pkl") learn.predict("Oliver Twist")</code></pre> <p><strong>Limitations and bias</strong></p> <p>The model was developed based on data from the&nbsp;<a href="https://www.bl.uk/collection-guides/digitised-printed-books">British Library&#39;s Digitised printed books (18th-19th Century)</a>&nbsp;collection. This dataset is not representative of books from the period covered with biases towards certain types (travel) and a likely absence of books that were difficult to digitise.</p> <p>The formatting of the British Library books corpus titles may differ from other collections, resulting in worse performance on other collections. It is recommended to evaluate the performance of the model before applying it to your own data. Likely, this model won&#39;t perform well for contemporary book titles without further fine-tuning.</p> <p><strong>Training data</strong></p> <p>The training data for this model will be available from the British Libary Research Repository shortly.</p> <p>The training data was created using the Zooniverse platform. British Library cataloguers carried out the majority of the annotations used as training data. More information on the process of creating the training data will be available soon.&nbsp;</p> <p><strong>Training procedure</strong></p> <p>Model training was carried out using the fastai library version 2.5.2.&nbsp;</p> <p>The notebook using for training the model will be available at: https://github.com/Living-with-machines/bl-books-genre-prediction</p> <p><strong>Eval result</strong></p> <p>The model was evaluated on a held out test set:</p> <pre> precision recall f1-score support Fiction 0.91 0.88 0.90 296 Non-fiction 0.94 0.95 0.95 554 accuracy 0.93 850 macro avg 0.93 0.92 0.92 850 weighted avg 0.93 0.93 0.93 850</pre>

openmit-licenseSep 2021View details →
zenodo44/100

Development of a spectral library for the discovery of altered genomic events in Mycobacterium avium associated with virulence using mass spectrometry-based proteogenomic analysis

<p><em>Mycobacterium avium</em> is one of the prominent disease-causing bacteria in humans. It causes lymphadenitis, chronic and extrapulmonary, and disseminated infections in adults, children, and immunocompromised patients. <em>M. avium</em> has ~4,500 predicted protein-coding regions on an average, which can be helpful in discovering several variants at the proteome level. Many of them are potentially associated with virulence, thus identifying such proteins can be a helpful feature in the development of panel-based theranostics. In line with such a long-term goal, we carried out an in-depth proteomic analysis of <em>M. avium</em> with both data-dependent and data-independent acquisition methods. Further, a set of proteogenomic investigations were carried out using the protein database for <em>Mycobacterium tuberculosis,</em> and a genome six-frame translated database and a variant protein database of <em>M. avium</em>. A search of mass spectrometry data analysis against <em>M. avium</em> protein database resulted in the identification of 2,954 proteins. Further, proteogenomic analyses aided in the identification of 1,301 novel peptide sequences and correction of translation start sites for 15 proteins. At the end, we created a spectral library of <em>M. avium</em> proteins including novel genome search-specific peptides and variant peptides detected in this study. We validated the spectral library by a data-independent acquisition of the <em>M. avium</em> proteome. Thus, we present a <em>M. avium </em>spectral library of 29,033 peptide precursors supported by 0.4 million fragment ions for further use by the biomedical community.</p>

opencc-by-4.0Nov 2022View details →
zenodo44/100

ACM-Digital Library (ACM)

<p>ACM-Digital Library (ACM): a subset of the ACM Digital Library with 24, 897 documents containing articles related to Computer Science. We considered only the first level of the taxonomy adopted by ACM, where each document is assigned to one of 11 classes.</p> <p>The files:<br> texts.txt: Document set (text). One per line.<br> score.txt: Document class whose index is associated with texts.txt<br> split_&lt;k&gt;.pkl:&nbsp;&nbsp;pandas DataFrame with k-cross validation partition.</p> <p>The .zip contains all aforementioned files + the tfidf representation in the CSR matrix format.</p>

opencc-by-4.0Aug 2021View details →
zenodo44/100

CoCO2 WP4.1 Library of Plumes

<p>This data collects synthetic CO2M images (the &#39;Library of Plumes&#39; for CoCO2 WP4.1) in the form of NetCDF files, for 5 modelling systems applied to 7 case studies. The files contain &#39;raw&#39; results (e.g., &quot;CO2_PP_M&quot;) in units of mol/m2, total column results (e.g., &quot;XCO2_PP_M&quot;) in units of ppm, and additionally the column mass (&quot;dry_mol_mass&quot; and surface pressure (&quot;surface pressure&quot;).</p> <p>This work was carried out in the context of the EU project CoCO2, by the following institutes: Empa (Switzerland), Deutsche Wetterdienst (Germany), Le Laboratoire des Sciences du Climat et de l&#39;Environnement (France), TNO (The Netherlands), Wageningen University &amp; Research (the Netherlands). More details about the model runs, their input data, possible issues, expected data quality, etc.,&nbsp;can be found in the accompanying report CoCO2 D4.2,&nbsp;<a href="https://www.coco2-project.eu/node/357">https://www.coco2-project.eu/node/357</a>.</p>

opencc-by-4.0Dec 2022View details →
zenodo44/100

Diachronic and diatopic word embeddings from newspapers digitised by the British Library (1830-1889): North and South England

<p>Diachronic&nbsp;word embeddings (decade-level) trained with Word2Vec (via Gensim) on different geographic subcorpora of the Heritage Made Digital British and the Living with Machines&nbsp;historical newspaper collections:</p> <p>- North England (north.zip)</p> <p>- South England (south.zip)</p> <p>At the moment, for each subcorpus, Word2Vec models are available for each decade in the period 1830-1889. More models are on the way for the following:</p> <p>- each decade in the periods 1780-1829 and 1890-1920 for both North and South England.</p> <p>- diachronic models for the following regions: Scotland, Wales, and Midlands.</p> <p>The models were trained using the following parameters:</p> <pre><code>sg = True min_count = 1 window = 5 vector_size = 200 epochs = 5</code></pre> <p>Like the embeddings in&nbsp;<a href="https://doi.org/10.5281/zenodo.7181682">this repository</a>, the model for each decade was aligned to the most recent one with Orthogonal Procrustes.</p> <p>See related GitHub repository for the full documentation:&nbsp;<a href="https://github.com/Living-with-machines/DiachronicEmb-BigHistData">https://github.com/Living-with-machines/DiachronicEmb-BigHistData</a>.</p> <p>Project website (Living with Machines):&nbsp;<a href="https://livingwithmachines.ac.uk/">https://livingwithmachines.ac.uk/</a></p> <p>Data related to:&nbsp;Nilo Pedrazzini &amp; Barbara McGillivray,&nbsp;<em>Diachronic and diatopic word embeddings from British historical newspapers</em>, presented at&nbsp;AIUCD (Convegno dell&rsquo;Associazione per l&rsquo;Informatica Umanistica e la Cultura Digitale) in Siena (Italy), June 2023.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

Spectral Library of European Pegmatites, Pegmatite Minerals and Pegmatite Host-Rocks – The Greenpeg Database

<p>Spectral signature, obtained through reflectance spectroscopy studies, of European pegmatites and minerals, as well of their host rocks. Samples include LCT- and NYF-type pegmatites and host rocks from pegmatite locations in Austria, Ireland, Norway, Portugal, and Spain. Sample preparation and spectral measurement were conducted in the Universidade do Porto &ndash; Faculdade de Ci&ecirc;ncias (UPORTO) laboratories. The database contains the reflectance spectra (raw and with continuum removed), sample photographs, and main absorption features automatically extracted by a Python routine. Whenever possible, spectral mineralogy was interpreted based on the continuum-removed spectra. A detailed description of the database, its content, the measuring instrument, and interoperability with GIS is found in the database report.</p>

opencc-by-4.0May 2022View details →
zenodo44/100

[ARAMCO Dhahran library collection 14430 AMK] ٱلْمَمْلَكَة ٱلْعَرَبِيَّة ٱلسُّعُوْدِيَّة. ٱلْهُفُوف

<p>[ARAMCO 14430 AMK] Kingdom of Saudi Arabia, al-Hufūf. Oblique aerial photograph of town and al-Kut, before 1973.</p>

opencc-by-4.0May 2023View details →
zenodo44/100

[ARAMCO Dhahran library collection 9051 TFW] ٱلْمَمْلَكَة ٱلْعَرَبِيَّة ٱلسُّعُوْدِيَّة. ٱلْهُفُوف .الكوت. قصرالأمير

<p>[ARAMCO 9051 TFW]&nbsp;Kingdom of Saudi Arabia, al-Hufūf. Emir&#39;s palace, as documented in the 1960s, now demolished. The current أمارة محافظة الأحساء is standing on the site.&nbsp;Coordinates: &nbsp;25&deg;22&#39;39&quot;N&nbsp;49&deg;35&#39;13&quot;E.</p>

opencc-by-4.0May 2023View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record