Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,093

datasets available to search

ShareScore release 0.7.1

Reset

Dataset results

1,093 results for “scripts”

Learn how ShareScore rates datasets ↗
zenodo52/100

Dataset and program scripts for the reproducibility of the hierarchical data structure file. Related to the manuscript entitled: Hierarchical Representation of Measurement Data, Metrological Uncertainty and Metadata for Calibrated Battery Tests

<p>We present an interoperable hierarchical data representation for battery tests, leading to improved scalability of data transmission and enhanced data accessibility and comprehensibility for both human interpretation and machine processing. The hierarchical data format includes the raw trace electrical measurement data, the metrological calibration and uncertainty data, the metadata such as experimental settings, instruments and software versions, as well as post-processed data such as electrochemical model fit parameters. This data representation allows repetition of the battery test under the exact same conditions such that identical results are achieved within defined error bounds. This is in line with the general F.A.I.R. data approach and provides repeatability and traceability in the battery value chain. As an application of the hierarchical data representation, we show the classification of cells as pass/fail being performed with quantitative confidence levels. We demonstrate the complete workflow of establishing the hierarchical data structure for electrochemical impedance spectroscopy (EIS), starting from metrological traceability of the calibration and uncertainty analysis towards the storage of the structured data as a single integrated file that preserves the hierarchical data format.</p>

openmit-licenseNov 2023View details →
zenodo52/100

Scripts for Patton et al 2020; Science DOI: 10.1126/science.abb9772

<p>Scripts used for all data anaysis for&nbsp; Patton et al. 2020 <em>Science&nbsp;</em>3<span>70: eabb9772. </span><span>DOI: 10.1126/science.abb9772</span></p>

opencc-by-4.0Nov 2024View details →
zenodo48/100

Dataset and Scripts for: RefPlantNLR: a comprehensive collection of experimentally validated plant NLRs (v.20200528_415)

<p><strong>RefPlantNLR v.20200528_415</strong></p> <p><strong>See&nbsp;</strong>bioRxiv&nbsp;2020.07.08.193961;&nbsp;doi:&nbsp;<a href="https://doi.org/10.1101/2020.07.08.193961">https://doi.org/10.1101/2020.07.08.193961</a></p> <p>SUPPLEMENTAL DATA</p> <p>Table S1: Description of RefPlantNLR.</p> <p>Table S2: Plant orders represented in RefPlantNLR.</p> <p>Supplemental dataset 1: Amino acid sequences of RefPlantNLR entries (fasta format). This file contains 415 amino acid sequences.</p> <p>Supplemental dataset 2: CDS sequences of RefPlantNLR entries (fasta format). This file contains 400 CDS sequences. CDS sequences could not be retrieved for 15 RefPlantNLR entries.</p> <p>Supplemental dataset 3: Annotated genomic sequences of RefPlantNLR entries (GenBank flat file format). This file contains 329 genomic loci containing the gene models of 344 RefPlantNLR entries and 56 RefPlantNLR mRNA entries lacking genomic information.</p> <p>Supplemental dataset 4: InterProScan annotation of the RefPlantNLR amino acid sequences (GFF3 format). This file contains the InterProScan annotation of 415 amino acid sequences.</p> <p>Supplemental dataset 5: InterProScan annotation of the RefPlantNLR CDS sequences (GFF3 format). This file contains the InterProScan annotation of the 400 CDS sequences.</p> <p>Supplemental dataset 6: Amino acid sequences of the extracted RefPlantNLR NB-ARC domains (fasta format). This file contains 424 NB-ARC domain (SUPERFAMILY signature SSF52540) amino acid sequences belonging to 415 RefPlantNLR entries.</p> <p>Supplemental dataset 7: Amino acid sequences of the unique RefPlantNLR extracted NB-ARC domains (fasta format). This file contains 347 unique NB-ARC domain (SUPERFAMILY signature SSF52540) amino acid sequences.</p> <p>Supplemental dataset 8: Clustal Omega alignment of the unique RefPlantNLR extracted NB-ARC domains (PHYLIP format). This file contains the Clustal Omega alignment of 346 unique NB-ARC domains (SUPERFAMILY signature SSF52540) with all positions with less than 95% coverage removed. Pb1 was omitted from this alignment.</p> <p>Supplemental dataset 9: NB-ARC domain phylogeny of the RefPlantNLR entries using the Maximum likelihood method (Newick format). This file contains the phylogenetic analysis of the NB-ARC domain of the RefPlantNLR entries using the JTT method.</p> <p>Supplemental dataset 10: Amino acid sequences of the non-redundant RefPlantNLR entries (fasta format). This file contains 235 amino acid sequences representing the non-redundant RefPlantNLR entries at a 90% amino acid identity threshold per genus according to the NB-ARC domain.</p> <p>Supplemental dataset 11: Amino acid sequences of the NB-ARC domains of the non-redundant RefPlantNLR entries (fasta format). This file contains 241 amino acid sequences representing the extracted NB-ARC domains of the 235 non-redundant RefPlantNLR.</p> <p>Appendix S1: R script used to generate annotations and figures.</p> <p>Appendix S2: InterProScan descriptions used for generating annotations.</p>

opencc-by-4.0Jul 2020View details →
zenodo48/100

Datasets and R-scripts used for the revision of the Dibrachys cavus complex by Peters & Baur, 2011, Zootaxa 2937.1

<p>In 2011 we published a revision on the Dibrachys cavus complex in Zootaxa (Peters and Baur, 2011, here a link to our <a href="https://doi.org/10.11646/zootaxa.2937.1.1">open access paper</a>). It was our wish to also publish two versions of the dataset (one with missing values, one with missing values imputed) as supplementary files. Unfortunately, the data files seem to be no longer available on the publishers webpage. Hence, we publish the data files herewith again in CSV format.</p> <p>We take the opportunity to also publish the R-scripts that we used for calculating multivariate analyses, tests, and the multiple imputation of missing values.</p> <p>All files are available individually and with an own link. For convenience, we have compiled all files also in a ZIP file.</p> <p><a href="https://doi.org/10.5281/zenodo.4256704">Baur (2020)</a> used the dataset for further exploration in a Multivariate Ratio Analysis (MRA).</p> <p>Papers quoted above you may find in the section <em>References</em> of the Zenodo package.</p> <p><strong>Citation of this package</strong><br> Peters, Ralph S., &amp; Baur, Hannes (2020, November 9) Datasets and R-scripts used for the revision of the Dibrachys cavus complex by Peters &amp; Baur, 2011, Zootaxa 2937.1. Zenodo. https://doi.org/10.5281/zenodo.4264539 (directs to the newest version of the package).</p>

opencc-by-4.0Nov 2020View details →
zenodo48/100

Code, data and scripts to study wave dynamics in asymmetric material

<p>Supplementary data for [1] Vladislav A. Yastrebov. &quot;Wave propagation through an elastically-asymmetric architected material&quot;, 2021, https://arxiv.org/abs/1712.06294v2</p> <p>See &quot;Readme.md&quot; and indivual &quot;Readme.md&quot; files in folders: &quot;data&quot;, &quot;src&quot;, &quot;fig&quot;</p>

opencc-by-4.0Jan 2021View details →
zenodo48/100

Data and script for the GenABEL paper

<p>This dataset contains the automatically collected data used for the overview paper about the GenABEL Project (Karssen et al, 2016, DOI:10.12688/f1000research.8733.1). Some data used for the paper was collected manually and is therefore not included in this dataset.</p> <p>The file &quot;tracker_report-2016-04-16.csv&quot; is an export of the bug reports from the GenABEL R-forge bug tracker on the date listed in the file name.</p> <p>The file &quot;Analytics www.genabel.org Locatie Lennart 20150428-20160428.csv&quot; is a custom export of the Google Analytics data for visits to the GenABEL website (www.genabel.org) in the period marked by the dates listed in the file name. The columns contain the ISO code of the country, city, number of sessions, number of new viewers, bounce percentage, pages per session and average session duration, respectively.</p> <p>The file analysis_GenABELpaper.org contains the source code<br> used for the automated data extraction for this paper in Emacs<br> Org mode literate programming format (http://orgmode.org, Schulte 2012, doi:10.18637/jss.v046.i03)</p>

opencc-by-sa-4.0May 2016View details →
zenodo48/100

Datafile DF1 + Matlab Scripts

<table> <tbody> <tr> <td>Supporting datafile&nbsp;</td> <td>DF1.csv</td> </tr> </tbody> </table> <table> <tbody> <tr> <td> <p>This file provides all experimental data associated with paper:&nbsp; Matrix gas flow through &lsquo;impermeable&rsquo; rocks - shales and tight sandstone , by Ernest Rutter, Julian Mecklenburgh and Yusuf Bashir</p> <p>Matlab Scripts is a compressed archive containing scripts for (a) the finite element simulaton of the stress state in a hydrostatically loaded sample with end piston constraints and (b) scripts for processing permeability data.&nbsp; A readme file is provided to explain the diferent fiels that make up the archive.</p> </td> </tr> </tbody> </table>

opencc-by-4.0Nov 2021View details →
zenodo48/100

Data and script for: "Increased birth rank of homosexual males: disentangling the older brother effect and sexual antagonism hypothesis"

<p>Data and script for Tables 2, 3, S2, S3, S4, and S5, and Figures 1, 3, 4, and S1.</p> <p>Individual dataset:</p> <p>France: <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/France_data_df12.csv">France_data_df12.csv </a><br> Indonesia: <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/Indonesia_data.csv">Indonesia_data.csv </a><br> Greece: <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/Greek_data.csv">Greek_data.csv </a></p> <p>The file <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/France_script.Rmd">France_script.Rmd </a>contains all the analyses of the french data set, including values presented Tables 2, 3, S3, S4, S5, Figures 3, 4 (output in file <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/France_script.html">France_script.html</a>). Same thing for files&nbsp;<a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/Indonesia_script.Rmd">Indonesia_script.Rmd&nbsp;</a> and <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/Greek_script.Rmd">Greek_script.Rmd</a>.</p> <p>For figure 1: <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/Fig1.html">Fig1.html </a><br> For Figure S1: <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/Fig_S1.html">Fig_S1.html </a><br> For Table S2: <a href="https://zenodo.org/api/files/acdb78f8-397c-4bc3-8f68-ec048edbc5f5/Table_S2_script.html">Table_S2_script.html </a><br> &nbsp;</p> <p>&nbsp;</p> <p><br> &nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Modified WRF/Chem source code, output data, and post-processing scripts for the GMD manuscript "Evaluation of WRF/Chem model (v3.9.1.1) real-time air quality forecasts over the Eastern Mediterranean"

<p>Here you will find the modified WRF/Chem code used in the simulations, the scripts used for post-processing and the model output data used in the manuscript.&nbsp;</p> <p>Two modifications have been made in&nbsp;module_aerosols_soa_vbs.F:</p> <ol> <li>ch_dust&nbsp;is set to1.0D-9*0.36</li> <li>The model is set not to initialize during restarts</li> </ol> <p>The model data directory includes:</p> <ol> <li>Two csv files (Winter and Summer) with the hourly concentrations of atmospheric pollutants&nbsp;at the locations of the ground stations. These data were used to produce Figures 4-8 in the manuscript as well as all the metrics.</li> <li>Two netcdf files&nbsp;(Winter and Summer) with the average ground concentrations of atmospheric pollutants over Cyprus. These data were use to produce Figure 3 in the manuscript.&nbsp;</li> </ol>

opencc-by-4.0Mar 2022View details →
zenodo48/100

Data, scripts, and R Notebook for Carneiro et al 2023. Flight performance and wing morphology in the bat Carollia perspicillata: biophysical models and energetics. Integrative Zoology DOI:10.1111/1749-4877.12707

<p>Files provided as supporting information for the paper by Carneiro et al. 2023. Flight performance and wing morphology in the bat&nbsp;<em>Carollia perspicillata</em>: biophysical models and energetics. Integrative Zoology. DOI:10.1111/1749-4877.12707</p> <p>File descriptions</p> <p>ArmTA.txt - Temperature and surface areas for arms of <em>C. perspicillata</em> after flight experiment<br> BodyTA.txt - Temperature and surface areas for body of <em>C. perspicillata</em> after flight experiment<br> HeadTA.txt - Temperature and surface areas for head of <em>C. perspicillata</em> after flight experiment<br> WingTA.txt - Temperature and surface areas for wings (patagium) of <em>C. perspicillata</em> after flight experiment<br> WingMorph.txt - Morphological variables measured in the body and wings of <em>C. perspicillata</em><br> HeatLoss.R - Function to estimate heat loss (Qt)<br> PowFlight.R - Function to estimate minimum power required to fly<br> Script-HeatLoss-FlightPerformance.R - R script with set of analyses performed<br> SupportingInformationFile.docx - R notebook with set of analyses performed, word format<br> SupportingInformationFile.nb.html - R notebook with set of analyses performed, html format<br> SupportingInformationFile.Rmd - R notebook with set of analyses performed (R markdown)</p> <p>For the R scripts (Script-HeatLoss-FlightPerformance.R) and notebook (<br> SupportingInformationFile.Rmd) to work and be compiled, all files need to be copied to the same folder.</p>

opencc-by-4.0Sep 2022View details →
zenodo48/100

Data and script for "On the emergence of ecosystem decay: a critical assessment of patch area effects across spatial scales"

<p>Data and R script necessary to replicate the results of Riva et al. 2024 ("On the emergence of ecosystem decay: a critical assessment of patch area effects across spatial scales"; minor revisions, Biological Conservation).</p>

opencc-by-4.0May 2024View details →
zenodo48/100

Data and analysis script for "The (non)effect of personalization in climate texts on credibility of climate scientists: A case study on sustainable travel"

<p>Dataset and analysis script for the article "<strong>The (non)effect of personalization in climate texts on credibility of climate scientists</strong><strong>: A case study on sustainable travel</strong>", under review at Geoscience Communication (https://doi.org/10.5194/egusphere-2024-543)</p>

opencc-by-4.0Jun 2024View details →
zenodo48/100

Data and script for Van Berkel et al: Can starlings use a reliable cue of future food deprivation to adaptively modify foraging and fat reserves?

<p>Supporting materials for:</p> <p><strong>Can starlings use a reliable cue of future food deprivation to adaptively modify foraging and fat reserves?</strong></p> <p>Menno van Berkel<sup>a</sup>, Melissa Bateson<sup>a</sup>, Daniel Nettle<sup>a</sup> and Jonathon Dunn<sup>a</sup>*</p> <p><sup>a</sup>Centre for Behaviour and Evolution &amp; Institute of Neuroscience, Newcastle University, Newcastle, UK</p> <p>*Author for correspondence (email: jonathon.dunn@newcastle.ac.uk; telephone: (+44)7730015855; postal address: Institute of Neuroscience, Henry Wellcome Building, The Medical School, Framlington Place, Newcastle University, Newcastle upon Tyne, UK, NE2 4HH).</p> <p>R script and 3 .csv files.</p>

opencc-by-4.0Mar 2018View details →
zenodo48/100

R script and data files for Oakley et al (2017) Journal of Proteome Research. DOI: 10.1021/acs.jproteome.6b00797

<p>This R script and data&nbsp;replicates the analysis&nbsp;of Oakley et&nbsp;al&nbsp;(2017) Thermal shock induces host proteostasis disruption and endoplasmic reticulum stress in the model symbiotic Cnidarian <em>Aiptasia</em>. <em>Journal of Proteome Research</em>. 16:2121-2134. DOI: 10.1021/acs.jproteome.6b00797.&nbsp;</p>

opencc-by-4.0Aug 2018View details →
zenodo48/100

Mitigating Network Noise on Dragonfly Networks through Application-Aware Routing (code, data and scripts to reproduce paper results)

<p>This repository contains the data, code, and scripts required to reproduce the results of the paper &quot;Mitigating Network Noise on Dragonfly Networks through Application-Aware Routing&quot; by Daniele De Sensi, Salvatore Di Girolamo and Torsten Hoefler, presented at the 2019 International Conference for High Performance Computing, Networking, Storage, and Analysis.&nbsp;</p> <p>This repository does not contains the code of the library used to automatically tune the routing algorithm, which can be found at http://doi.org/10.5281/zenodo.3372785</p>

opencc-by-4.0Aug 2019View details →
zenodo48/100

Dataset and R script for the analysis in the article "Food waste between environmental education, peers, and family influence. Insights from primary school students in Northern Italy", Journal of Cleaner Production

<p>We hereby publish the dataset (with metadata) and the R script (R Core team, 2018) used for implementing the analysis presented in the paper&nbsp;&quot;Food waste between environmental education, peers, and family influence. Insights from primary school students in Northern Italy&quot;,&nbsp;<em>Journal of Cleaner Production </em>(Piras et al., 2023). The dataset is provided in csv format with semicolons as separators and &quot;NA&quot; for missing data. The dataset&nbsp;includes all the variables used in at least one of the models presented in the paper, either in the main text or in&nbsp;the Supplementary Material. Other variables gathered by means of the questionnaires included as Supplementary Material of the paper have been removed. The dataset includes inputted values&nbsp;for missing data on independent variables. These were inputted using two approaches: last observation carried forward (LOCF) - preferred when possible -&nbsp;and last observation carried backward (LOCB). The metadata are presented as a PDF file.</p>

opencc-by-4.0Nov 2022View details →
zenodo48/100

CA-discharge data set, scripts and raw data

<p>Gauge locations of 295 gauges in Central Asia, including long-term norm discharge, basin outlines and basin characteristics compiled from 3rd-party data. Discharge time series for 135 gauge locations collected from the hydrological yearbooks of Hydrometeorological Organisations in Central Asia.&nbsp;</p> <p>Instructions of how to use the data can be found in the readme document in the folder CA-data-paper-scripts.&nbsp;</p> <p>! Important note: Please do not use the glacier thinning rates extracted from Hugonnet et al., 2021 (https://doi.org/10.1038/s41586-021-03436-z), i.e. features gl_dmdt_km3a and gl_dmdtda_mma. The ice density is not accounted for in our averages.&nbsp;</p>

opencc-by-4.0Mar 2023View details →
zenodo44/100

Eyebright species maintenance (scripts and data accompanying Becher et al., Plant Communications)

<p><strong>This gzipped TAR ball contains data and scripts related to the study on Fair Isle eyebrights by Hannes Becher, Max R. Brown, Gavin Powell, Chris Metherell, Nick J. Riddiford, and Alex D. Twyford, submitted to Plant Communications.</strong></p> <p>Data: genome assembly of<em> Euphrasia arctica</em>, variant call files of the &quot;tetraploid&quot; and &quot;conserved&quot; sets of scaffolds, per-individual k-mer spectra, mapping depths, etc.</p> <p>Scripts: R scripts for the analysis of plant trait data, heterozygosity, ect.; an ipython notebook for the analysis of variant data, and a Mathematica notebook with the derivation of the formulae used to fit pop gen parameters to k-mer spectra.</p>

opencc-by-4.0Apr 2020View details →
zenodo44/100

Data and R-Scripts for "Quality and timing of crowd-based water level class observations"

<p>This are the data and the R-scripts used for the manuscript &quot;Quality and timing of crowd-based water level class observations&quot; accepted for publication in the journal Hydrological Processes in July 2020 as a Scientific Briefing. To run the code, just run the R-script with the name &quot;RunThisForResults.R&quot;. Results will be written to the &quot;Figures&quot; and the &quot;Results&quot; folder.</p>

opencc-by-4.0Feb 2020View details →
zenodo44/100

1QIsaa data collection (binarized images, feature files, and plotting scripts) for writer identification test using artificial intelligence and image-based pattern recognition techniques

<p><strong>The Great Isaiah Scroll (1QIsa<sup>a</sup>) data set for writer identification</strong></p> <p>This data set is collected for the ERC project:<br> The Hands that Wrote the Bible: Digital Palaeography and Scribal Culture of the Dead Sea Scrolls<br> PI: Mladen Popović<br> Grant agreement ID: 640497</p> <p>Project website: <a href="https://cordis.europa.eu/project/id/640497">https://cordis.europa.eu/project/id/640497</a><br> <br> <strong>Copyright (c) </strong>&nbsp;&nbsp; &nbsp;University of Groningen, 2021. All rights reserved.<br> <strong>Disclaimer and copyright notice for all data contained on this .tar.gz file:</strong></p> <p><strong>1)</strong> permission is hereby granted to use the data for research purposes. It is not allowed to distribute this data for commercial purposes.</p> <p><strong>2) </strong>provider gives no express or implied warranty of any kind, and any implied warranties of merchantability and fitness for purpose are disclaimed.</p> <p><strong>3) </strong>provider shall not be liable for any direct, indirect, special, incidental, or consequential damages arising out of any use of this data.</p> <p><strong>4) </strong>the user should refer to the first public article on this data set:<br> <br> <em>Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence-based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.</em><br> <br> BibTeX:</p> <pre>@article{popovic2020artificial, title={Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsaa)}, author={Popovi{\&#39;c}, Mladen and Dhali, Maruf A and Schomaker, Lambert}, journal={arXiv preprint arXiv:2010.14476}, year={2020} }</pre> <p><strong>5) </strong>the recipient should refrain from proliferating the data set to third parties external to his/her local research group. Please refer interested researchers to this site for obtaining their own copy.</p> <p><strong>Organisation of the data:</strong></p> <p>The .tar.gz file contains three directories: images, features, and plots. The included &#39;README&#39; file contains all the instructions.</p> <p>The &#39;images&#39; directory contains NetPBM images of the columns of 1QIsa<sup>a</sup>. The NetPBM format is chosen because of its simplicity. Additionally, there is no doubt about lossy compression in the processing chain. There are two images for each of the Great Isaiah Scroll columns: one is the direct binarized output from the BiNet (<em>arxiv.org/abs/1911.07930</em>) system, and the other one is the manually cleaned version of the binarized output. &nbsp; The file names for the direct binarized output are of the format &#39;1QIsaa_col&lt;columnnr&gt;.pbm&#39;, for example, &#39;1QIsaa_col15.pbm&#39;. And, for the cleaned version, the format is &#39;1QIsaa_col&lt;columnnr&gt;_cleaned.pbm&#39;, for example, &#39;1QIsaa_col15_cleaned.pbm&#39;. Note: the image files are not in a separate directory; they will be extracted in the same place. However, due to the unique naming, there is no problem extracting them in one single directory.</p> <p>The &#39;features&#39; directory contains feature files computed for each of the column images. There are two types of feature files: Hinge and Adjoined. They are distinguishable by their extension, for example, &#39;1QIsaa_col15_cleaned.hinge&#39; and &#39;1QIsaa_col15_cleaned.adjoined&#39;. They are also arranged in separate directories for ease of use.</p> <p>The &#39;plots&#39; directory contains a simple python script to perform PCA on the feature files and then visualize them in a 3D plot. The file takes the location of feature files as an input. The &#39;README_plot&#39; file contains examples of how-to-run in the terminal.</p> <p><strong>Brief description:</strong><br> According to ImageMagick&#39;s&#39; identify&#39; tool, the original images are in grayscale (.jpg) from Brill collection, in &#39;8-bit Gray 256c&#39;. &nbsp;These images pass through multiple preprocessing measures to become suitable for pattern recognition-based techniques. The first step in preprocessing is the image-binarization technique. In order to prevent any classification of the text-column images based on irrelevant background patterns, a specific binarization technique (BiNet) was applied, keeping the original ink traces intact. After performing the binarization, the images were cleaned further by removing the adjacent columns that partially appear on the target columns&#39; images. Finally, few minor affine transformations and stretching corrections were performed in a restrictive manner. These corrections are also targeted for aligning the texts where the text lines get twisted due to the leather writing surface&#39;s degradation. Hence, the clean images are there in the directory along with the direct binarized images. No effort has been made to obtain a balanced set in any way.</p> <p><strong>Tools:</strong><br> <strong>Binarization:</strong><br> The BiNet tool is available for scientific use upon request (m.a.dhal(at)rug.nl)</p> <p><strong>Image Morphing:</strong><br> In the original article, data augmentation was performed using image morphing. The tool is available on GitHub:<br> https://github.com/GrHound/imagemorph.c</p> <p><strong>Features for writer identification:</strong><br> Lambert Schomaker<br> http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> <em><strong>1.&nbsp;</strong>L. Schomaker &amp; M. Bulacu (2004). Automatic writer identification using connected-component contours and edge-based features of upper-case Western script. IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.<br> <strong>2. </strong>Bulacu, M. &amp; Schomaker, L.R.B. (2007). Text-independent Writer Identification and Verification Using Textural and Allographic Features, &nbsp;IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</em><br> &nbsp;<br> The features (hinge, fraglets) have been combined in a single MS Windows application, GIWIS, which is available for scientific use upon request (l.r.b.schomaker(at)rug.nl)</p> <p><strong>If you have any question, please contact us:</strong><br> Maruf A. Dhali &lt;m.a.dhali(at)rug.nl&gt;<br> Lambert Schomaker &lt;l.r.b.schomaker(at)rug.nl&gt;<br> Mladen Popović &lt;m.popovic(at)rug.nl&gt;</p> <p><strong>Please cite our papers if you use this data set:</strong><br> <em><strong>1.</strong> Popović, M., Dhali, M. A., &amp; Schomaker, L. (2020). Artificial intelligence based writer identification generates new evidence for the unknown scribes of the Dead Sea Scrolls exemplified by the Great Isaiah Scroll (1QIsa<sup>a</sup>). arXiv preprint arXiv:2010.14476.<br> <strong>2. </strong>Dhali, M. A., de Wit, J. W., &amp; Schomaker, L. (2019). Binet: Degraded-manuscript binarization in diverse document textures and layouts using deep encoder-decoder networks. arXiv preprint arXiv:1911.07930.</em></p>

opencc-by-4.0Jan 2021View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record