Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
249
datasets available to search
ShareScore release 0.9.0
Dataset results
249 results for “letter”
Chemical data accompanying the manuscript "Chromium cycling in redox-stratified basins challenges δ53Cr paleoredox proxy applications" in Geophysical research Letters
<p>Water column and sediment chromium concentration and stable isotope data and ancillary metal data from Lake Cadagno, Switzerland. These data accompany a manuscript by the same authors in Geophysical Research Letters (doi: 10.1029/2022GL099154).</p> <p> </p> <p>The associated CTD data are available in the following Zenodo dataset: Sepúlveda Steiner, O., Carlino, C., Haizmann, E., Roman, S., Wüest, A., & Bouffard, D. (2022). Lake Cadagno 2017 CTD and water quality monitoring [Data set]. Zenodo. <a href="http://doi.org/10.5281/zenodo.7127882">http://doi.org/10.5281/zenodo.7127882</a></p>
Mississippi River spatial water chemistry Environmental Research Letters datasets
We mapped surface water chemistry along the entire length of the Upper Mississippi River (UMR) to understand spatial patterns in nitrate sources and processing. We used a sensor-based and boat-mounted sensing platform to continuously measure underway water chemistry. Measurements were linked with global positioning systems (GPS) to create maps of surface water chemistry. Here, we archive data associated with an Environmental Research Letters publication (Loken et al. 2018). Data include a single spatial survey of the entire length of the UMR (Minneapolis, Minnesota to Cairo, Illinois) in August 2015 and repeat surveys in Navigation Pool 8 (located near La Crosse, WI). Data have been provided in three formats (raw, hydraulic-corrected, and tau-corrected). Additionally, we archive laboratory chemistry data from water samples collected during the project. Sites include a range of main channel, backwaters, and tributaries. Water chemistry samples were analyzed at the North Temperate Lakes - Long Term Ecological Research facility and linked with underway sensor measurements.
Photosynthetic quotients in aquatic ecosystems: data and code supporting Trentman et al. 2023 manuscript in L&O Letters
This study provides a summary of the mismatch between our current knowledge and the application of the photosynthetic quotient (PQ). We use data from the Upper Clark Fork River (UCFR) as a case study example of how the PQ may vary in space and time based on environmental conditions. Surface water sample measurements of dissolved oxygen (DO), temperature (T), nutrients (NO3-N, NH4-N, SRP), and several metabolism indicators are represented in this data product. Figures represent data from two sites on the mainstem of the Upper Clark Fork River (UCFR) over a roughly two-year period, from 2019 to 2021. Some measurements are derived from existing data products or manuscripts, including DOT (Valett, et al., 2023); nutrients (H. M. Valett, Dec. 2, 2022, pers. comm); air pressure (Deer Lodge Weather Station, 2023); underlying data for Trentman et al. (2023) Figure 2 and Figure 4e and 4f (via Burris, 1981); and SI-Figure2 USGS gage data (USGS, 2023). Products unique to this data product include metabolism data (Trentman, et al., 2023 (Figure 5)), chamber data supporting Trentman, et al., (2023) Figure 6, and code simulations/data. All analytes and variables are documented in the project data dictionary. For details on data collection methods, see the methods section, the manuscript, and/or referenced data products.
MCR LTER: Coral Reef: Spatial portfolios in coral metapopulations are shaped by spatiotemporal asynchrony in environmental conditions; Data for Srednick et al., 2026 Ecology Letters
Using wavelet analyses of a 19-year coral community timeseries from Moorea, French Polynesia, we quantified timescale-specific population synchrony in four common coral genera and evaluated the predictors of spatial portfolio effects. We detected synchrony within genera associated with synchrony in degree heating days, diurnal temperature range (DTR), and macroalgal cover at different timescales. Synchrony in DTR and macroalgal cover was associated with lower synchrony of Pocillopora and Porites populations, respectively. Population (for three of four genera) and environmental synchrony were stronger within than among habitats across timescales, underscoring the role of habitat-specific conditions in driving spatial synchrony and spatial portfolios. These results describe how the spatial and temporal scales of heterogeneity in environmental and ecological conditions determine synchrony in coral population dynamics and support a spatial portfolio effect, which may buffer coral metapopulations from island-scale collapse. Data in support of analyses for: Spatial portfolios in coral metapopulations are shaped by spatiotemporal asynchrony in environmental conditions. Published in Ecology Letters 2026.
Dataset supporting the paper "Doublet-Singlet-Doublet Transition in a Single Organic Molecule Magnet On-Surface Constructed with up to 3 Aluminum Atoms. Nano Letters 21, 8317 (2021)"
<p>Dataset corresponding to theoretical calculations in the paper "Doublet-Singlet-Doublet Transition in a Single Organic Molecule Magnet On-Surface Constructed with up to 3 Aluminum Atoms" Nano Letters 21, 8317 (2021), <a href="https://doi.org/10.1021/acs.nanolett.1c02881">https://doi.org/10.1021/acs.nanolett.1c02881</a></p> <p>List of files:</p> <p>Several folders corresponding to the figures of the paper. They contain:</p> <ul> <li>.siesta files: STM images in WsXM format (http://www.wsxm.eu/) simulated using STMpw (<a href="https://doi.org/10.5281/zenodo.3581159">https://doi.org/10.5281/zenodo.3581159</a>).</li> <li>CONTCAR and POSCAR files: relaxed structures in VASP format. They can be visualized with VESTA (<a href="https://jp-minerals.org/vesta/en/">https://jp-minerals.org/vesta/en/</a>).</li> <li>.agr: grace files (<a href="https://plasma-gate.weizmann.ac.il/Grace/">https://plasma-gate.weizmann.ac.il/Grace/</a>).<br> </li> </ul>
Typewriter indentations in a letter by W. H. Auden - visualized with photometric stereo
<p>This figure shows a section of a letter from W. H. Auden to Stella Musulin.</p> <p>Top: conventional photograph.<br> Middle: raking light photograph. Indentations in the paper become visible, but hardly legible.<br> Bottom: a false-color visualization generated with photometric stereo. The simultaneous display of depth, albedo and curvature gradient allow the discrimination of four text layers.</p>
Data from: Padfield et al. (2016) Rapid evolution of metabolic traits explains thermal adaptation in phytoplankton. Ecology letters.
<p>This repository provides the data from the TPC and logistic growth curves from the paper:</p> <p>Padfield, D., Yvon‐Durocher, G., Buckling, A., Jennings, S., & Yvon‐Durocher, G. (2016). Rapid evolution of metabolic traits explains thermal adaptation in phytoplankton. Ecology letters, 19(2), 133-142.</p> <p>metadata.pdf gives a more detailed explanation of the data.</p>
"ICDAR2023 Competition on Detection and Recognition of Greek Letters on Papyri" Dataset
<h1><strong>Dataset description of the “ICDAR2023 Competition on Detection and Recognition of Greek Letters on Papyri”</strong></h1> <p>Prof. Dr. Isabelle Marthot-Santaniello, Dr. Olga Serbaeva</p> <p>2024.09.16</p> <h2>Introduction</h2> <p>The present dataset stems from the ICDAR2023 Competition on Detection and Recognition of Greek Letters on Papyri (original links to the competition are provided in the file “1b.CompetitionLinks.”)</p> <p>The aim of this competition was to investigate the performance of glyph detection and recognition in a very challenging type of historical document: Greek papyri. The detection and recognition of Greek letters on papyri is a preliminary step for computational analysis of handwriting that can lead to major steps forward in our understanding of this important source of information on Antiquity. Such detection and recognition can be done manually by trained papyrologists. It is, however, a time-consuming task that would need automatising. </p> <p>We provide here the documents related to two different tasks: localisation and classification. The document images are provided by several institutions and are representative of the diversity of book hands on papyri (a millennium time span, various script styles, provenance, states of preservation, means of digitization and resolution).</p> <h2>How the dataset was constructed</h2> <p>In the frame of <a href="https://d-scribes.philhist.unibas.ch/en/case-studies/iliad-208/" target="_blank" rel="noopener">D-Scribes project</a> lead by Prof. Dr. Isabelle Marthot-Santaniello, 2018-2023, around 150 papyri fragments containing Iliad were manually annotated at a letter-level in <a href="https://github.com/readsoftware/read" target="_blank" rel="noopener">READ</a>.</p> <p>The editions were taken, for the major part, from <a href="papyri.info" target="_blank" rel="noopener">papyri.info</a>, and were simplified, i.e. the accents, editorial marks, and other additional information were removed to be as close as possible to what is to be found on papyri. When the text was not available on papyri.info, the relevant passage was extracted from the <a href="https://github.com/PerseusDL/canonical-greekLit/blob/master/data/tlg0012/tlg001/tlg0012.tlg001.perseus-grc2.xml" target="_blank" rel="noopener">Homer Iliad of Perseus</a>.</p> <p>From those, 150 plus papyri fragments, 185 surfaces (sides of fragments) belonging to 136 different manuscript identified by their Trismegistos numbers, (further TMs) were selected to serve as a material for Competition. These 185 surfaces were separated into the “training set” and the “test set” provided for the competition as a set of images and corresponding data in JSON format.</p> <p>Details on the competition summarised in "ICDAR 2023 Competition on Detection and Recognition of Greek Letters on Papyri", by Mathias Seuret, Isabelle Marthot-Santaniello, Stephen A. White, Olga Serbaeva Saraogi, Selaudin Agolli, Guillaume Carrière, Dalia Rodriguez-Salas, and Vincent Christlein; edited by G. A. Fink et al. (Eds.): <em>ICDAR 2023,</em> LNCS 14188, pp. 498–507, 2023. https://doi.org/10.1007/978-3-031-41679-8_29.</p> <p>After the competition ended, the decision was taken to release manually annotated dataset for the “test set” as well. Please find the description of each included document below.</p> <h2><br>Dataset Structure</h2> <p><br><strong>“1. CompetitionOverview.xlsx”</strong> contains the metadata of the used images in Excel file, state 2024.09.19. Here is the structure of the Excel file:</p> <p> </p> <table> <tbody> <tr> <td> <p><strong>Excel columns</strong></p> </td> <td> <p><strong>Name</strong></p> </td> <td> <p><strong>Content</strong></p> </td> <td> <p><strong>Notes</strong></p> </td> </tr> <tr> <td> <p><strong>A</strong></p> </td> <td> <p><strong>TM</strong></p> </td> <td> <p><strong>Trismegistos number is internationally used for papyri identification</strong></p> </td> <td> <p><strong>With READ item name in ().</strong></p> </td> </tr> <tr> <td> <p><strong>B</strong></p> </td> <td> <p><strong>Papyri.info link</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p> </p> </td> </tr> <tr> <td> <p><strong>C</strong></p> </td> <td> <p><strong>Fragments' Owning Institution (from <a href="http://papyri.info"><u>papyri.info</u></a>) </strong></p> </td> <td> <p><strong>Institution’s name</strong></p> </td> <td> <p><strong>Institution that physically stores the papyri</strong></p> </td> </tr> <tr> <td> <p><strong>D</strong></p> </td> <td> <p><strong>Availability (of metadata, <a href="http://papyri.info"><u>papyri.info</u></a>) </strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>Metadata reuse clarification</strong></p> </td> </tr> <tr> <td> <p><strong>E</strong></p> </td> <td> <p><strong>text ID (READ)</strong></p> </td> <td> <p><strong>Number from READ SQL database that was used to link the images and the editions.</strong></p> </td> <td> <p><strong>Serves to locate the attached images and understand the JSON structure.</strong></p> </td> </tr> <tr> <td> <p><strong>F</strong></p> </td> <td> <p><strong>Test/Training</strong></p> </td> <td> <p> </p> </td> <td> <p><strong> I.e. the image was originally included in the training or in the test set of the dataset.</strong></p> </td> </tr> <tr> <td> <p><strong>G</strong></p> </td> <td> <p><strong>Image Name (for orientation)</strong></p> </td> <td> <p> </p> </td> <td> <p><strong>As in READ</strong></p> </td> </tr> <tr> <td> <p><strong>H</strong></p> </td> <td> <p><strong>Cedopal link</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>Contains additional metadata and includes the links to all available online images.</strong></p> </td> </tr> <tr> <td> <p><strong>I</strong></p> </td> <td> <p><strong>License from the Institution webpage.</strong></p> </td> <td> <p><strong>Either license or usage summary.</strong></p> </td> <td> <p><strong>If no precise licence has been given, the summary of the reuse rights is provided with a link to the regulations in column K</strong></p> </td> </tr> <tr> <td> <p><strong>J</strong></p> </td> <td> <p><strong>Image URL</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>Not all images are available online. Please contact the owning institution directly if the image is not available.</strong></p> </td> </tr> <tr> <td> <p><strong>K</strong></p> </td> <td> <p><strong>Information on the image usage from the institution</strong></p> </td> <td> <p><strong>link</strong></p> </td> <td> <p><strong>In case of any doubt, please contact the owning institution directly.</strong></p> </td> </tr> <tr> <td> <p><strong>L</strong></p> </td> <td> <p><strong>Notes</strong></p> </td> <td> <p> </p> </td> <td> <p> </p> </td> </tr> </tbody> </table> <p>For the purpose of an easy overview, the items with special problems, i.e. images not online or missing links, have been marked in red.</p> <p><strong>2. There are three data subsets:</strong></p> <p><strong>2a. “Training file” </strong><br>(containing 150 papyri images separated into 108 texts and HomerCompTraining.json). The images are those of papyri containing Iliad of Homer in JPG-format. These were processed in READ, namely, each visible letter on a given papyri was linked to the edition of the Iliad, through this process, each linked letter of the edition was linked to its coordinates in pixels on the HTML-surface of the image. All that information is provided in the JSON-file.</p> <p>The JSON file contains the <strong>“annotations”</strong> (b-boxes of each letter/sign), <strong>“categories”</strong> (Greek letters),<strong> “images”</strong> (Image IDs), and <strong>“licenses”</strong>. The links between image and bboxes is defined via the “id” in the “images” part (for example, "id": 6109). This same id is encoded as “"image_id": 6109” in the “annotations”. Alternatively, “text_id” which can be found in the “images” URL and in the file-names provided here and containing images, can be used for data linking.</p> <p>Let us now describe the content of each part of the JSON file:<br>Each <strong>“annotation”</strong> contains<br>“area" characterised as “bbox" with coordinates, <br>“category_id”, that allows to identify which Greek letter in categories is represented by the number; “id”, which is a unique number of the cliplet, i.e. area; <br>“image_id”, that links cliplet to the surface of the image having the same id; <br>“iscrowd" and “seg_id" are useful to find the information back in READ database; <br>and, finally, “tags”.</p> <p>In tags, “BaseType" was used to annotate quality as described below. “FootMarkType”, ft1, etc., was used for clustering tests, but played no role for the Competition.<br>“BaseType” ot bt-tags were assigned to the letters to mark the quality of preservation: <br>bt-1: well-preserved letter that should allows easy identification for both human eyes and the Computer-vision; <br>bt-2: Partially preserved letter that might also have some background damage (holes, additional ink, etc), but remains readable, and has one interpretation. <br>bt-3: Letters damaged to such an extant that they cannot be identified without reading an edition. These are treated as traces of ink. <br>bt-4: The letters that have some damage, but this damage is of such kind that it makes possible multiple interpretations. For example, missing/defaced horizontal stroke makes alpha indistinguishable from damaged delta or lambda.</p> <p>Each <strong>“category”</strong> contains <br>“id”, this is a number references also in “annotations” and it allows to identify which Greek letter was in the bbox; <br>”name”, for example, “χ”; <br>and “supercategory”, i.e. “Greek”.</p> <p>Each <strong>“image”</strong> contains the following sub fields: <br>“bln_id" is an internal READ number of the html surface; <br>"date_captured": null - is another READ field; <br>"file_name": “./images/homer2/txt1/P.Corn.Inv.MSS.A.101.XIII.jpg", allows to link easy image and text, i.e. for the image in question the JPG will be in the file called “txt1”, it is very similar by structure and function to "img_url": "./images/homer2/txt1/P.Corn.Inv.MSS.A.101.XIII.jpg"; <br>each image has “height" and “width" expressed in pixels. <br>Each image has “id”, and this id is referenced in the “annotations” under “image_id”. <br>Finally, each image contains a link to “license”, expressed as a number. </p> <p>Each <strong>“licence”</strong> lists a license as it was found during the time of competition, i.e. in February 2023.</p> <p><strong>2b. “Test file”</strong> <br>contains 34 papyri image sides separated into 31 TMs and HomerCompTesting.json The JSON file here only allows to connect the images with the “categories”, “images”, “licenses”, but without the “annotations”. The structure and logic is otherwise the same like in “Training” JSON.</p> <p><strong>2c. “Answers file” </strong><br>Containing the “annotations” and other information for the 34 papyri of the “Testing” dataset. The structure and logic is the same like in “Training” JSON.</p> <p><strong>3. “Additional files” </strong><br>Containing lists of duplicate segments id (multiple possible readings or tags), respectively 6 items for “Training”, 17 for “Testing” and 15 for “Answers”.</p> <p><strong>4. “Dataset Description”</strong><br>This same description included for completeness.</p> <h2>References</h2> <p>The Dataset was reused or mentioned in a number of publications (state September 2024)</p> <p>Mohammed, H., Jampour, M. (2024). "From Detection to Modelling: An End-to-End Paleographic System for Analysing Historical Handwriting Styles". In: Sfikas, G., Retsinas, G. (eds) <em>Document Analysis Systems. DAS 2024.</em> Lecture Notes in Computer Science, vol 14994. Springer, Cham, pp. 363–376. https://doi.org/10.1007/978-3-031-70442-0_22</p> <p>De Gregorio, G., Perrin, S., Pena, R.C.G., Marthot-Santaniello, I., Mouchère, H. (2024). "NeuroPapyri: A Deep Attention Embedding Network for Handwritten Papyri Retrieval". In: Mouchère, H., Zhu, A. (eds) <em>Document Analysis and Recognition – ICDAR 2024 Workshops. ICDAR 2024.</em> Lecture Notes in Computer Science, vol 14936. Springer, Cham, pp. 71–86. https://doi.org/10.1007/978-3-031-70642-4_5</p> <div> <p>Vu, M. T., Beurton-Aimar, M. "PapyTwin net: a Twin network for Greek letters detection on ancient Papyri". <em>HIP '23: 7th International Workshop on Historical Document Imaging and Processing, San Jose, CA, USA, August 2023.</em><br>https://doi.org/10.1145/3604951.3605522<br>https://dl.acm.org/doi/fullHtml/10.1145/3604951.3605522</p> <p>Turnbull, R., Mannix, E. "Detecting and recognizing characters in Greek papyri with YOLOv8, DeiT and SimCLR". (Preprint).<br>arXiv:2401.12513<br>https://doi.org/10.48550/arXiv.2401.12513</p> </div>
Dataset supporting the paper "Thioetherification of Br-Mercaptobiphenyl Molecules on Au(111). Nano Letters 23, 1350 (2023)"
<p>Dataset corresponding to theoretical calculations in the paper "Thioetherification of Br-Mercaptobiphenyl Molecules on Au(111). Nano Letters 23, 1350 (2023)" DOI: <a href="https://doi.org/10.1021/acs.nanolett.2c04619">https://doi.org/10.1021/acs.nanolett.2c04619</a></p> <p>List of files:</p> <p>Several folders corresponding to the figures of the paper. They contain:</p> <ul> <li>.dat files: STM images simulated using STMpw (<a href="https://doi.org/10.5281/zenodo.3581159">https://doi.org/10.5281/zenodo.3581159</a>). They can be processed with the programs and scripts in the Utils directory of STMpw.</li> <li>CONTCAR files: relaxed structures in VASP format. They can be visualized with VESTA (<a href="https://jp-minerals.org/vesta/en/">https://jp-minerals.org/vesta/en/</a>).</li> <li>.agr: grace files (<a href="https://plasma-gate.weizmann.ac.il/Grace/">https://plasma-gate.weizmann.ac.il/Grace/</a>).</li> </ul>
Numbers and Letters
Open the record for dataset details and reuse information.
Data used to create figures in the ACP Letters manuscipt "The value of remote marine aerosol measurements for constraining radiative forcing uncertainty" by Regayre et al. (2020)
<p>This dataset was created from perturbed parameter ensembles (PPEs) using the HadGEM-UKCA atmospheric composition climate model. All data needed to reproduce figures in the Regayre et al. (2020) ACP Letters article "The value of remote marine aerosol measurements for constraining radiative forcing uncertainty" are included. Other output from the PPEs can be obtained by contacting the lead author.</p> <p>The following data are included here:</p> <ul> <li>CCN measurement data degraded to match the model-measurement comparison resolution.</li> <li>Unconstrained and constrained CCN<sub>0.2</sub> output from the PPE used to make Figure 1. These compressed files contain 48 .dat files. Each .dat file contains the PPE mean, variance and 95% creidble interval data. Files are named consecutively, containing data from 90<sup>o</sup>S to 90<sup>o</sup>N at 0<sup>o</sup>E, then continuing Eastward. When combined, these files provide data for each latitude/longitude pair at the N48 spatial resolution.</li> <li>A zip file of an netcdf file containing 26-dimensional data for parameter values, used to create the sample of 1 million model variants from our statistical emulators of model output.</li> <li>A zip file containing a folder of files made of one million ones and zeros that indicate the retention/rejection criteria from applying our constraint methodology for various constraint combination scenarios, for each model variant. A value of 1 indicates the model variant was retained. Data in these files is in the same order as the unconstrained sample file of parameter values.</li> <li>Compressed files containing global, annual mean RF<sub>aci</sub> and ERF<sub>aci</sub> values for the unconstrained set of one million model variants. The compressed netcdf files contain RF (ERF), RF<sub>aci</sub> (ERF<sub>aci</sub>) and RF<sub>ari</sub> (ERF<sub>ari</sub>) values.</li> </ul>
A Survey of Body Part Construction Metaphors in the Neo-Assyrian Letter Corpus
<p>The dataset consists of approximately 2,400 examples of metaphors in Akkadian of what we term Body Part Constructions (BPC's) within the letter sub-corpus of the <a href="http://oracc.museum.upenn.edu/saao/">State Archives of Assyria online</a> (SAAo). The dataset was generated by a multi-step process involving the training and application of a spaCy language model to the SAAo letter sub-corpus, converting the resulting annotations to linked open data format amenable to searching for BPC’s, and manually adding metalinguistic data to the search results; these files, in CONLLU and TTL formats, as well as the model specific files based on spaCy's requirements, are also made available in this publication. The BPC dataset is stored as a CSV file, and can serve as an easy starting place for other scholars interested in finding socio-linguistic usage patterns of this construction.</p> <p>The royal archives of the late Neo-Assyrian kings (8th-7th century BCE) constitute an important source for understanding many facets of the Neo-Assyrian empire. Ranging from treaty tablets and legal documents to prophecies, ritual instructions, and even court literature, the approximately five thousand texts in this corpus primarily come from the palatial complex at Nineveh and document the reigns of Sargon II (r. 721-705), Sennacherib (r. 704-681), Esarhaddon (r. 680-669), and Assurbanipal (668-627). Over the past four decades, much of these archives has been published in the State Archives of Assyria (SAA) volumes at the University of Helsinki, and in more recent years has appeared digitally under the <a href="http://www.en.ag.geschichte.uni-muenchen.de/research/mocci/">Munich Open-access Cuneiform Corpus Initiative</a> (LMU Munich) as the SAAo.</p>
Data set for letter "Floquet-Driven Crossover from Density-Assisted Tunneling to Enhanced Pair Tunneling"
<p>The files contain the data depicted in the figures of the article "Floquet-Driven Crossover from Density-Assisted Tunneling to Enhanced Pair Tunneling", arXiv 2404.08482.</p> <p>The format of the data and to which figure it corresponds is described in the file "README.txt".</p>
Tactile Braille Letters Dataset
<p>Contains reading of single braille letters by sliding a robotic tactile sensor over each letter of the braille alphabet, including space. Braille letters are part of the dataset in the STL file format. Tactile data is pulled with 40Hz and saved in the pandas format and dumped in the pickle format. Sigma-delta modulation was used to create event-driven (spike-based) binary discrete-time data streams with four different encoding threshold. For more details please see the linked <a href="https://github.com/event-driven-robotics/tactile_braille_reading">Github repository</a>.</p>
Frauen* im Fokus. Transcriptions and full texts of letters and works of women's rights activists
<p>In winter 2023/24 the Berlin State Library and Potsdam University (chair for Comparative Literature) organized the citizen science workshop <a href="https://lab.sbb.berlin/events/frauen-im-fokus/">"Frauen* im Fokus"</a> (women* in focus).</p> <p>Within the project 48 participants transcribed 85 letters and documents by 19th- and early 20th-century women's rights activists held in the collections of the State Library. The transcriptions provided by the participants were aligned with the digital images of the items using the <a href="https://ocr-bw.bib.uni-mannheim.de/escriptorium/" target="_blank" rel="noopener">eScriptorium</a> platform and software and checked for potential errors. From eScriptorium, the transcriptions were exported as ALTO and PAGE files. From the PAGE files, TEI/XML files were created with added metadata about the correspondence (where applicable).</p> <p>In addition to these MS sources, 55 printed works of the same activists were OCRed and are included in the dataset as PAGE and ALTO-files; these files were not manually checked for quality.</p> <p>We would like to thank our trainees Lilly Bucksteeg and Lilly Welz for their valuable contribution to the creation of the data set.</p> <p>The data set consists of five zip-files containing:</p> <ul> <li>the print sources in PAGE format</li> <li>the print sources in ALTO format</li> <li>the manuscript sources in PAGE format</li> <li>the manuscript sources in ALTO format</li> <li>the manuscript sources in TEI format</li> </ul> <p>------------------</p> <p>Im Wintersemester 2023/24 führten die Staatsbibliothek zu Berlin und die Universität Potsdam (Professur für Allgemeine und Vergleichende Literaturwissenschaft) das <a href="https://lab.sbb.berlin/events/frauen-im-fokus/">Projekt "Frauen* im Fokus"</a> durch.</p> <p>Im Rahmen des Projekts transkribierten 48 Personen 85 Briefe und andere Nachlassdokumente von Frauenrechtlerinnen des 19. und frühen 20. Jahrhunderts. Das Organisationsteam nahm auf der Plattform <a href="https://ocr-bw.bib.uni-mannheim.de/escriptorium/" target="_blank" rel="noopener">eScriptorium</a> eine teilweise automatisierte Layoutanalyse und Zeilensegmentierung der Digitalisate vor und fügte nach erfolgter inhaltlicher Qualitätskontrolle die von den Teilnehmenden erstellten Transkriptionen dort ein, um sie dann als PAGE- und ALTO-Dateien zu exportieren; zusätzlich wurden aus den PAGE-Dateien TEI-Dateien der einzelnen Dokumente generiert, die (wo passend) mit Brief-Metadaten zu Absendern, Empfängern und Orten angereichert wurden.</p> <p>Im Rahmen des Projekts wurden zudem 55 Druckwerke der Frauenrechtlerinnen aus dem Bestand der Staatsbibliothek als Volltexte erschlossen. Die Segmentierung und Volltexterkennung der Druckwerke erfolgte automatisch in eScriptorium ohne zusätzliche manuelle Qualitätskontrolle. Die Daten liegen exportiert in den Formaten PAGE und ALTO vor.</p> <p>Besonderer Dank gebührt Lilly Bucksteeg und Lilly Welz, die als Praktiantinnen im Projekt maßgeblich zur Erstellung des Datensets beigetragen haben.</p> <p>Das Datenset besteht aus fünf zip-Dateien die folgende Dateien enthalten:</p> <ul> <li>die gedruckten Werke im PAGE-Format</li> <li>die gedruckten Werke im ALTO-Format</li> <li>die Manuskript-Transkriptionen im PAGE-Format</li> <li>die Manuskript-Transkriptionen im ALTO-Format</li> <li>die Manuskript-Transkriptionen im TEI-Format</li> </ul>
Derived data and analysis code accompanying Deines et al. 2019, Environmental Research Letters
<p>This codebase accompanies the paper:</p> <p>Deines, JM, AD Kendall, JJ Butler, Jr., & DW Hyndman. 2019. Quantifying irrigation adaptation strategies in response to stakeholder-driven groundwater management in the US High Plains Aquifer. Environmental Research Letters. DOI: <a href="https://doi.org/10.1088/1748-9326/aafe39">https://doi.org/10.1088/1748-9326/aafe39</a></p> <p>Data and code at time of publication.</p>
Edward FitzGerald Life and Letters – Overview of all letters
<p>These files form part of an archive of research material relating to the Life and Letters of Edward FitzGerald. The data have been compiled by independent researchers W H (Bill) Martin and Sandra Mason; their contact details are below. The files comprise a number of searchable listings of information contained in FitzGerald’s letters. The information formed input to a book on Edward FitzGerald which is referenced below.</p> <p>This section of the archive contains a large database, <em>efgdb</em>, giving an overview of the letters, their dating, the people to whom FitzGerald wrote, his location at the time of writing, and a broad classification of the content of each letter. The letters are those contained in the collection published by A M Terhune and A B Terhune in 1980 – see reference below. An explanatory README text file contains further information on the database, including a table showing the fields included in the database and giving definitions of them and of the codings used where relevant. </p>
Edward FitzGerald Life and Letters – Analysis of writing and reading
<p>These files form part of an archive of research material relating to the Life and Letters of Edward FitzGerald. The data have been compiled by independent researchers W H (Bill) Martin and Sandra Mason; their contact details are below. The files comprise a number of searchable listings of information contained in FitzGerald’s letters. The information formed input to a book on Edward FitzGerald which is referenced below.</p> <p>This section of the archive contains a database, <em>efgliterature</em>, which analyses the letters in terms of their comments on FitzGerald’s writings and the books and other material that he read. It also shows the dating of the letters, the people to whom FitzGerald wrote, and his location at the time of writing. The letters are those contained in the collection published by A M Terhune and A B Terhune in 1980 – see reference below. The database contains all the letters published by the Terhunes, including some which have no literary references. An explanatory README text file contains a table showing the fields included in the database and giving definitions of them and of the codings used where relevant. </p> <p> </p>
Edward FitzGerald Life and Letters – Analysis of other interests and views
<p>These files form part of an archive of research material relating to the Life and Letters of Edward FitzGerald. The data have been compiled by independent researchers W H (Bill) Martin and Sandra Mason; their contact details are below. The files comprise a number of searchable listings of information contained in FitzGerald’s letters. The information formed input to a book on Edward FitzGerald which is referenced below.</p> <p>This section of the archive contains a number of databases which analyse the letters in terms of their comments on FitzGerald’s interests and activities, and his views on a variety of topics. The databases cover current affairs (<em>efgctaffairs</em>), religion (<em>efgreligion</em>), travel (<em>efgtravel</em>), leisure activities (<em>efgactivities</em>), nature and countryside (<em>efgnature</em>), food and drink (<em>efgfood</em>), and personal matters (<em>efgcharacter</em>). They also show the dating of the letters, the people to whom FitzGerald wrote, and his location at the time of writing. The letters are those contained in the collection published by A M Terhune and A B Terhune in 1980 – see reference below. An explanatory README text file contains tables showing the fields included in the databases and giving definitions of them and of the codings used where relevant. </p>
Edward FitzGerald Life and Letters – Analysis of family and friends
<p>These files form part of an archive of research material relating to the Life and Letters of Edward FitzGerald. The data have been compiled by independent researchers W H (Bill) Martin and Sandra Mason; their contact details are below. The files comprise a number of searchable listings of information contained in FitzGerald’s letters. The information formed input to a book on Edward FitzGerald which is referenced below.</p> <p>This section of the archive contains a database, <em>efgpeople</em>, which analyses the letters in terms of their comments on FitzGerald’s family and a wide range of his friends. It also shows the dating of the letters, the people to whom FitzGerald wrote, and his location at the time of writing. The letters are those contained in the collection published by A M Terhune and A B Terhune in 1980 – see reference below. The database contains all the letters published by the Terhunes, including some which have no references to family and friends. An explanatory README text file contains a table showing the fields included in the database and giving definitions of them and of the codings used where relevant. </p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.