Skip to main content
Powered by ShareScore

Find research datasets worth reusing

Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.

1,549

datasets available to search

ShareScore release 0.9.0

Reset

Dataset results

1,549 results for “benchmarks”

Learn how ShareScore rates datasets ↗
zenodo44/100

Koios benchmarks' netlist files (BLIF format)

<p>Koios is a benchmark suite for FPGA architecture and CAD exploration. It contains circuits from the Deep Learning domain.</p> <p>Here we are uploading the netlist files for these designs obtained by synthesizing the Verilog designs using Intel Quartus for Intel Stratix IV&nbsp;architecture.&nbsp;</p> <p>For more details see:&nbsp;</p> <p>https://docs.verilogtorouting.org/en/latest/vtr/benchmarks/#koios-benchmarks</p> <p>https://github.com/verilog-to-routing/vtr-verilog-to-routing/tree/master/vtr_flow/benchmarks/verilog/koios</p>

opencc-by-4.0Oct 2022View details →
zenodo44/100

Benchmarking tools for transcription factor prioritization

<p><strong>Abstract:</strong></p> <p>Spatiotemporal regulation of gene expression is controlled by transcription factor (TF) binding to regulatory elements, resulting in a plethora of cell types and cell states from the same genetic information.&nbsp; Due to the importance of regulatory elements, various sequencing methods have been developed to localise them in genomes, for example using ChIP-seq profiling of the histone mark H3K27ac that marks active regulatory regions. Moreover, multiple tools have been developed to predict TF binding to these regulatory elements based on DNA sequence. As altered gene expression is a hallmark of disease phenotypes, identifying TFs driving such gene expression programs is critical for the identification of novel drug targets.In this study, we curated 84 chromatin profiling experiments (H3K27ac ChIP-seq) where TFs were perturbed through e.g., genetic knockout or overexpression. We ran nine published tools to prioritize TFs using these real-world data sets and evaluated the performance of the methods in identifying the perturbed TFs. This allowed the nomination of three frontrunner tools, namely RcisTarget, MEIRLOP and monaLisa. Our analyses revealed opportunities and commonalities of tools that will help to guide further improvements and developments in the field.</p> <p><strong>Dataset description:</strong></p> <ul> <li>tf_tool_benchmark_atacseq_diffPeaks.tar.gz -Archive containing differential peak statistics, tool diff peak input files (fore- and background) for all currated ATAC-seq datasets.&nbsp;</li> <li>tf_tool_benchmark_h3K27ac_chipseq_diffPeaks.tar.gz - Archive containing differential peak statistics, tool diff peak input files (fore- and background) for all currated H3K27ac ChIP-seq datasets.&nbsp;</li> <li>tf_tool_benchmark_atacseq_results.tar.gz - Archive containing the raw tool results for each ATAC-seq dataset.</li> <li>tf_tool_benchmark_chipseq_results.tar.gz - Archive containing the raw tool results for each H3K27ac ChIP-seq dataset.</li> <li>tf_tool_benchmark_results.tar.gz - Archive containing tool results summary for plotting (rds files).</li> </ul> <p><strong>Contact:&nbsp; </strong>Sebastian Steinhauser - sebastian.steinhauser@novartis.com</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

FG-OVD: Fine-grained Open-Vocabulary Object Detection Benchmark Suite

<p>A collection of annotations for PACO images containing free-form fine-grained textual captions of objects, their parts, and their attributes. It also comprises several sets of negative captions that can be used to test and evaluate the fine-grained recognition ability of open-vocabulary models.</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

Joint AstraZeneca-Cancer Research Horizons Functional Genomics Centre's CRISPRn library benchmark screens: gRNA counts and associated metadata

<p>Genome-wide CRISPR sgRNA libraries have emerged as transformative tools to systematically probe gene function. While these libraries have been iterated over time to be more efficient, their large size limits their use in some applications. Here, we benchmarked publicly available genome-wide single-targeting sgRNA libraries and evaluated dual targeting as a strategy for pooled CRISPR loss-of-function screens. We leveraged this data to design two minimal genome-wide human CRISPR-Cas9 libraries that are 50% smaller than other libraries and that preserve specificity and sensitivity, thus enabling broader deployment at scale.&nbsp;</p>

opencc-by-4.0May 2024View details →
zenodo44/100

FLOW-Alaiz Benchmark: Baseline Results

<p>This repository hosts documentation and notebooks for the <a href="https://www.flow-horizon.eu/">EU-FLOW</a> Alaiz benchmark.</p> <p>The first stage is finished with the publication of baseline results for PyWAsP, PyWAsP-CFD, and SiteFlow in the Torque-2024 conference.</p> <p><strong>Sanz Rodrigo J, Oxley G and Tobias Olsen B (2024) FLOW-Alaiz benchmark for coupled terrain and array interaction flow models. Baseline Results. J. Phys.: Conf. Ser. 2767 092077, <a href="https://iopscience.iop.org/article/10.1088/1742-6596/2767/9/092077">doi:10.1088/1742-6596/2767/9/092077</a></strong></p>

opencc-by-4.0Jun 2024View details →
zenodo44/100

Bloume: Evaluation Results for French Benchmarks

<p>These datasets contain the outputs of evaluation results for the BLOOM large language model tasked with a set of standard NLP benchmarks in French. The project is fully documented on github. See: https://github.com/fyvo/EvaluerBloom.&nbsp;</p>

opencc-by-4.0Dec 2023View details →
zenodo44/100

Canonical Decision Diagrams Modulo Theories - Benchmarking

<p>This archive contains all data and results that were used to benchmark the approach Canonical Decision Diagrams Modulo Theories. The archive contains the following files: 3 subfolders, one for each dataset that was used in the benchmarking process, each one containing a "data" subfolder which contains the problems (in SMT/SMT2 format) and some output folders which contain JSON files describing in detail the results of each run on the problems.</p>

opencc-by-4.0Jul 2024View details →
zenodo44/100

Firemaker image collection for benchmarking forensic writer identification using image-based pattern recognition

<p>Disclaimer and terms of use:<br> ============================</p> <p>/*****************************************************************************\<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; This is the Firemaker NFI-images Distribution &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; This distribution contains 1000 images of scanned handwritten text, &nbsp; &nbsp; &nbsp; *<br> * &nbsp; scanned at resolution 300dpi grey scale, containing pages of &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; handwritten text by 250 writers, four pages per writer, from four &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; writing conditions, one condition per page. The conditions are: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; p1: copied, natural style, p2: copied, UPPER case, p3: copied and forged, *<br> * &nbsp; i.e.,&quot;try to write in a different style than your natural style&quot;, and p4, *<br> * &nbsp; self generated, i.e., text produced to describe a given cartoon. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; Copyright The International Unipen Foundation, 2000, All rights reserved &nbsp;*<br> *******************************************************************************<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CDROM: &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; PURPOSES. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; DISCLAIMED. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL, &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; DATA. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;4) THE USER SHOULD REFER TO THE FIRST PUBLIC ARTICLE ON THIS DATA SET: &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; M. Bulacu, L. Schomaker &amp; L. Vuurpijl (2003). &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; Writer identification using edge-based directional features. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; ICDAR &#39;03: Proceedings of the 7th International Conference on Document &nbsp;*<br> * &nbsp; &nbsp; Analysis and Recognition, pp. 937-941. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;*<br> * &nbsp; &nbsp; Piscataway: IEEE Computer, ISBN 0-7695-1960-1 &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> * &nbsp;5) THE RECIPIENT SHOULD REFRAIN FROM PROLIFERATING THE DATA SET TO THIRD &nbsp; *<br> * &nbsp;PARTIES EXTERNAL TO HIS/HER LOCAL RESEARCH GROUP. PLEASE REFER INTERESTED &nbsp;*<br> * &nbsp;RESEARCHERS TO HTTP://UNIPEN.ORG FOR OBTAINING THEIR OWN COPY. &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; *<br> \*****************************************************************************/</p> <p>BibTeX entry: &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;&nbsp;</p> <p>&nbsp; @inproceedings{Firemaker, &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;<br> &nbsp; &nbsp; author = {Bulacu, M. and Schomaker, L.R.B. and Vuurpijl, L.}, &nbsp; &nbsp;<br> &nbsp; &nbsp; title = {Writer Identification Using Edge-Based Directional Features},<br> &nbsp; &nbsp; booktitle = {ICDAR &#39;03: Proceedings of the 7th International&nbsp;<br> &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp; &nbsp;Conference on Document Analysis and Recognition},<br> &nbsp; &nbsp; year = {2003},<br> &nbsp; &nbsp; isbn = {0-7695-1960-1},<br> &nbsp; &nbsp; pages = {937-941},<br> &nbsp; &nbsp; publisher = {IEEE Computer Society},<br> &nbsp; &nbsp; address = {Washington, DC, USA},<br> &nbsp; &nbsp;}</p> <p>In the project &quot;Vergelijk&quot;, a grant obtained from the Dutch Forensic Science<br> Institute, two existing professional writer-identification systems have been&nbsp;<br> compared regarding usability studies and in particular recognition&nbsp;<br> performance (Schomaker &amp; Vuurpijl, 2000). The results of this comparison&nbsp;<br> are contained in a confidential report:</p> <p>&nbsp; L.R.B. Schomaker and L.G. Vuurpijl (2000).&nbsp;<br> &nbsp; Forensic writer identification: A benchmark data set&nbsp;<br> &nbsp; and a comparison of two systems. Technical report,&nbsp;<br> &nbsp; Nijmegen Institute for Cognition and Information (NICI),&nbsp;<br> &nbsp; University of Nijmegen, The Netherlands.</p> <p>Informative and non-confidential details from this report are&nbsp;<br> given in the accompanying file: &nbsp;&#39;firemaker-dbase.pdf&#39;</p> <p>To compare both systems, a carefully designed experiment was conducted to<br> record handwritten samples from male and female writers in several conditions:</p> <p>Condition 1: Normal constrained handwriting<br> ==============================================</p> <p>Below, the Dutch text writers had to produce in normal handwriting is given.&nbsp;</p> <p>--- start text ----<br> Zij bezochten veilingen en reisden met de KLM. Voor<br> korte afstanden huurden ze een auto, meestal een VW<br> of een Ford.<br> &lt;EMPTY LINE&gt;<br> De veilingen waren van 7-4-1993 tot 3-5-1993 in New<br> York, Tokyo, Qu&eacute;bec, Rome, Parijs, Z&uuml;rich en Oslo.<br> &lt;EMPTY LINE&gt;<br> Omdat de veilingen steeds begonnen om 12 uur en je<br> gemiddeld 200 tot 300 kilometer moest rijden,<br> stonden zij steeds om 6.30 uur op en vertrokken om<br> 8 uur uit het hotel.<br> &lt;EMPTY LINE&gt;<br> Elke dag hadden ze vijfhonderd (f 500,-) gulden<br> nodig. Daarvoor gebruikten ze elke keer een cheque<br> van tweehonderd (f 200,-) en een cheque van<br> driehonderd (f 300,-) gulden. Aan geschenken gaven<br> ze ongeveer honderd gulden (f 100,-) uit.<br> --- end text ----</p> <p><br> Condition 2: Production of constrained block capital handwriting<br> ================================================================</p> <p>In this condition, the writers had to produce the following text<br> in block-capital handwriting:</p> <p>--- start text ----<br> NADAT ZE IN NEW YORK, TOKYO, QU&Eacute;BEC, PARIJS, Z&Uuml;RICH<br> EN OSLO WAREN GEWEEST, VLOGEN ZE UIT DE USA TERUG<br> MET VLUCHT KL 658 OM 12 UUR.<br> &lt;empty line&gt;<br> ZE KWAMEN AAN IN DUBLIN OM 7 UUR EN IN AMSTERDAM OM<br> 9.40 UUR &#39;S AVONDS. DE FIAT VAN BOB EN DE VW VAN<br> DAVID STONDEN IN R3 VAN HET PARKEERTERREIN.<br> HIERVOOR MOESTEN ZE HONDERD GULDEN (F 100,-)<br> BETALEN.<br> --- end text ----</p> <p><br> Condition 3: Production of free-forged handwriting<br> ==================================================</p> <p>Below, the text writers had to produce in the free-forged handwriting<br> condition is given. No example of handwriting is given which they have to<br> mimick (forge), the condition concerns a self-conceived distorted&nbsp;<br> handwriting style.</p> <p>--- start text ----<br> Nog dezelfde avond reden ze naar hun vrienden<br> Chris, Emile, Jan, Irene en Henk, nadat ze hun<br> vriendinnen Greta en Maria hadden opgehaald.<br> &lt;EMPTY LINE&gt;<br> Samen hadden ze vijfhonderd (500) zeldzame<br> postzegels gekocht, Bob driehonderd (300) en David<br> tweehonderd (200).<br> &lt;EMPTY LINE&gt;<br> De reis was de moeite waard geweest.<br> --- end text ----</p> <p><br> Condition 4: Production of unconstrained handwriting<br> ====================================================</p> <p>The final text writers had to produce is unconstrained handwriting.<br> The cartoon, a series of pictures concerning a &#39;UFO&#39; landing had<br> to be described in their own words, in at least six lines of text.<br> See image file &quot;space.gif&quot;.</p> <p><br> Thruth labels and writer identifications<br> ========================================</p> <p>Each writer has a unique id, specified as:</p> <p>&nbsp; &nbsp;id: &nbsp; {num}{set}<br> &nbsp; num: &nbsp; a three-digit number<br> &nbsp;set: &nbsp; &nbsp;either 01, 02, 03 or 04, identifying one of the 4 experiments</p> <p>The vast majority of the writers producing sets 01, 02 and 03 mimicked the<br> content and layout (empty lines) of the constrained texts they had to copy<br> sufficiently accurately, such that the example texts are a good indication of<br> the contents. However, as set 04 (&quot;describe cartoon story&quot;) &nbsp;contains<br> unconstrained self-generated handwriting, the corresponding thruth &nbsp;labels had<br> to be extracted manually. The resulting label files are contained in &nbsp;the<br> directory ./300dpi/p4-self-natural/labels/</p> <p>Note: no letter, word, line or paragraph segmentation is provided with this<br> data set. The main text can be cropped easily. Since the orientation is<br> horizontal, projection techniques can be used to extract lines, using<br> a line-spacing parameter (~94 pixels line height) as an additional check.&nbsp;</p> <p><br> Overview of directories:</p> <p>300dpi/<br> &nbsp; &nbsp;p1-copy-normal/ &nbsp; &nbsp; &nbsp;Copying task, normal writing style &nbsp;<br> &nbsp; &nbsp;p2-copy-upper/ &nbsp; &nbsp; &nbsp; Copying task, UPPER-case&nbsp;<br> &nbsp; &nbsp;p3-copy-forged/ &nbsp; &nbsp; &nbsp;Copying task, instructed to mimic another script style<br> &nbsp; &nbsp;p4-self-natural/ &nbsp; &nbsp; Self-generated text, natural writing condition</p> <p>Note: the original raw collection contained writer #155, who has been removed<br> from this data set, as his first condition (p1) was started in upper case and<br> the page was not &nbsp;completed. Deleted files were 15501.tif, 15502.tif, 15503.tif<br> and 15504.tif.</p> <p>Note: the name of this data set (Firemaker) is a contraction of the names<br> Vuurpijl and Schomaker.</p> <p>Note b: Example of a cutout of essential handwritten text using NetPBM tools: &nbsp;<br> &nbsp;tifftopnm 15201.tif | pnmcut -left 50 -right 2400 -top 700 -bottom 3250 &gt; handwriting.pgm</p> <p>&nbsp;For an experiment, the upper and lower halves of the resulting image were<br> &nbsp;usually used in the Schomaker &amp; Bulacu studies to obtain two samples of&nbsp;<br> &nbsp;handwriting for a writer.</p> <p>&nbsp;http://www.ai.rug.nl/~lambert<br> &nbsp;http://www.ai.rug.nl/~bulacu</p> <p>Our features for writer identification:</p> <p>Lambert Schomaker<br> &nbsp;http://www.ai.rug.nl/~lambert/allographic-fraglet-codebooks/allographic-fraglet-codebooks.html<br> &nbsp;L. Schomaker &amp; M. Bulacu (2004).&nbsp;<br> &nbsp;Automatic writer identification using connected-component contours and edge-based features of upper-case Western script.&nbsp;<br> &nbsp;IEEE Transactions on Pattern Analysis and Machine Intelligence, Vol 26(6), June 2004, pp. 787 - 798.</p> <p>Marius Bulacu<br> &nbsp;http://www.ai.rug.nl/~lambert/hinge/hinge-transform.html<br> &nbsp;Bulacu, M. &amp; Schomaker, L.R.B. (2007).&nbsp;<br> &nbsp;Text-independent Writer Identification and Verification Using Textural and Allographic Features,&nbsp;<br> &nbsp;IEEE Trans. on Pattern Analysis and Machine Intelligence (PAMI), Special Issue - Biometrics: Progress and Directions, April, 29(4), p. 701-717.</p> <p>Axel Brink<br> &nbsp;http://www.ai.rug.nl/~axel/ &nbsp;&#39;Quill&#39; feature<br> &nbsp;A.A. Brink, J. Smit, M.L. Bulacu, and L.R.B. Schomaker (2011).&nbsp;<br> &nbsp;Writer identification using directional ink-trace width measurements,&nbsp;<br> &nbsp;Pattern Recognition (July 2011), doi: 10.1016/j.patcog.2011.07.005<br> &nbsp;<br> These three feature groups (hinge, fraglets, quill) have been combined in<br> a single MS Windows application, GIWIS which is available for scientific<br> use upon request (schomaker@ai.rug.nl)</p> <p>Note c.</p> <p>The accompanying file &#39;Firemaker-writer-info.dat&#39; contains some<br> writer information:&nbsp;<br> Column 1: writer identification code<br> Column 2: sex<br> Column 3: handedness,&nbsp;<br> Column 4: age in years<br> Column 5: major Western script group (print,cursive or mixed)<br> &nbsp;</p>

opencc-by-4.0Dec 1999View details →
zenodo44/100

Data set and benchmarks from Eriksson et al., ICAPS 2018

<p>These are the benchmarks and the experiment data used in the paper &quot;A Proof System for Unsolvable Planning Tasks&quot; by Eriksson et al. (ICAPS 2018). The raw data contains the logs from all runs, while the eval-directories contain a json file with all parsed attributes. The benchmarks directory contains all benchmarks used in the experiments. Finally, the file eriksson-et-al-icaps2018.html provides an overview over the most interesting attributes across all runs.</p>

opencc-by-4.0Mar 2018View details →
zenodo44/100

TriGraphSlant - benchmark set for writer identification - writers were asked to write in unnatural slant

<p>&nbsp;<br> Disclaimer and terms of use:<br> ============================<br> &nbsp;<br> /*****************************************************************************\<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; This is the TrigraphSlant (Img version) Distribution, release 18/3/2011&nbsp;&nbsp; *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; This distribution contains 188 images of scanned handwritten text,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; scanned at resolution 300dpi Canon LiDE 25, grey scale,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; by 47 Dutch writers, four pages per writer, from four&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; writing conditions, one condition per page. The conditions are:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; 1. [AN] Copy text A in your natural handwriting.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; 2. [BN] Copy text B in your natural handwriting.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; 3. [BL] Copy text B and slant your handwriting to the&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; left as much as possible.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; 4. [BR] Copy text B and slant your handwriting to the&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; right as much as possible.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; The codes AN, BN, BL and BR refer to subsets into which the collected&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; pages of the writers were subdivided. AN represents a collection of&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; authentic documents; BN, BL and BR can be seen as collections of&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; questioned documents. To avoid structural effects of fatigue, the order&nbsp;&nbsp; *<br> *&nbsp;&nbsp; of item 3 and 4 was randomized at each collection: half of the subjects&nbsp;&nbsp; *<br> *&nbsp;&nbsp; wrote the BR page before the BL page. The data were collected at three&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; sites, in three cities: The Hague: NFI (N...), Donders Institute for&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; Brain, Cognition and Behaviour, Radboud University Nijmegen (D...)&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; and the Artificial Intelligence Dept. of University of Groningen (R...)&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp; Copyright The International Unipen Foundation, 2010, All rights reserved&nbsp; *<br> *******************************************************************************<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp; DISCLAIMER AND COPYRIGHT NOTICE FOR ALL DATA CONTAINED ON THIS CARRIER:&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp; 1) PERMISSION IS HEREBY GRANTED TO USE THE DATA FOR RESEARCH&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp; PURPOSES. IT IS NOT ALLOWED TO DISTRIBUTE THIS DATA FOR COMMERCIAL&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp; PURPOSES.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp; 2) PROVIDER GIVES NO EXPRESS OR IMPLIED WARRANTY OF ANY KIND AND ANY&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp; IMPLIED WARRANTIES OF MERCHANTABILITY AND FITNESS FOR PURPOSE ARE&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp; DISCLAIMED.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp; 3) PROVIDER SHALL NOT BE LIABLE FOR ANY DIRECT, INDIRECT, SPECIAL,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp; INCIDENTAL OR CONSEQUENTIAL DAMAGES ARISING OUT OF ANY USE OF THIS&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp; DATA.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp; 4) THE USER SHOULD REFER TO THE FOLLOWING ARTICLE ON THIS DATA SET:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> * A.A. Brink, R.M.J. Niels, R.A. van Batenburg, C.E. van den Heuvel,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> * L.R.B. Schomaker, Towards robust writer verification by correcting&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> * unnatural slant, Pattern Recognition Letters, Volume 32, Issue 3,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> * 1 February 2011, Pages 449-457, ISSN 0167-8655,&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> * DOI: 10.1016/j.patrec.2010.10.010.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> *&nbsp; 5) THE RECIPIENT SHOULD REFRAIN FROM PROLIFERATING THE DATA SET TO THIRD&nbsp;&nbsp; *<br> *&nbsp; PARTIES EXTERNAL TO HIS/HER LOCAL RESEARCH GROUP. PLEASE REFER INTERESTED&nbsp; *<br> *&nbsp; RESEARCHERS TO HTTP://UNIPEN.ORG FOR OBTAINING THEIR OWN COPY.&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; *<br> \*****************************************************************************/<br> &nbsp;<br> Abstract<br> &nbsp;<br> Towards robust writer verification by correcting unnatural slant<br> &nbsp;<br> A.A. Brink, , R.M.J. Niels, R.A. van Batenburg, C.E. van den Heuvel,&nbsp; &nbsp;<br> and L.R.B. Schomaker, &nbsp;<br> &nbsp;<br> a Institute of Artificial Intelligence and Cognitive Engineering (ALICE), &nbsp;<br> &nbsp; University of Groningen, P.O. Box 407, 9700 AK Groningen, The Netherlands<br> &nbsp; &nbsp;<br> b Donders Institute for Brain, Cognition and Behaviour, Radboud University Nijmegen, &nbsp;<br> &nbsp; P.O. Box 9104, 6500 HE Nijmegen, The Netherlands<br> &nbsp; &nbsp;<br> c Netherlands Forensic Institute, P.O. Box 24044, 2490 AA Den Haag, The Netherlands<br> &nbsp;<br> Received 11 September 2009.&nbsp; Available online 30 October 2010.<br> &nbsp;<br> Slant is a salient feature of Western handwriting and it is considered to be an<br> important writer-specific feature. In disguised handwriting however, slant is<br> often modified. It was tested whether slant is indeed an important factor and it<br> was tested whether the distorting effect of deliberate slant change can be<br> countered by a simple shear transform. This was done in two off-line writer<br> verification experiments in image processing conditions of slant elimination and<br> slant correction. The experiments were performed using three features based on<br> statistical pattern recognition, including the state-of-the-art features<br> Fraglets and Hinge. A new public dataset was created and used, containing<br> natural and slanted handwriting by 47 writers. A striking result is that the<br> average natural slant value is much less important for biometric systems than is<br> usually assumed: eliminating slant yields just a 1-5% performance loss. A<br> second result is that the effects of deliberate slant change cannot be fully<br> countered by a simple shear transform: it raises performance on the distorted<br> handwriting from 53-68% to 64-90%, but this is still lower than normal<br> operation on natural handwriting: 97-100%.<br> &nbsp;<br> Research highlights<br> - The value of slant as a writer identification feature has been overrated. &nbsp;<br> - Deliberate slant change can be partly countered by the shear transform. &nbsp;<br> - Deliberate slant change introduces non-affine distortions to the handwriting. &nbsp;<br> - A new dataset of deliberately slanted handwriting was introduced.<br> &nbsp;<br> Keywords: Handwriting biometrics; Writer verification; Slant; Disguise; Statistical<br> pattern recognition</p> <p>&nbsp;</p>

opencc-by-4.0Mar 2011View details →
zenodo44/100

Benchmark comparison tests between Ambit-SMIRKS and RDKit chemoinformatics tools

<p>This archive contains benchmark code and results for Ambit-SMIRKS software package (<a href="http://ambit.sf.net">http://ambit.sf.net</a>) , described in the publication &ldquo;Kochev N,, Avramova S., Jeliazkova N. Ambit-SMIRKS: a Software Module for Reaction Representation, Reaction Search and Structure Transformation&rdquo;.&nbsp;</p> <p>We have performed benchmark testing of Ambit-SMIRKS and RDKit SMIRKS transformation algorithms. For this purpose we used a set of 545 compounds (see file <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/smiles-set.txt?versionId=a49915af-23b7-41d2-9625-4a90cee172b4">smiles-set.txt</a>) including normal constituents of the body or common components of food, provided by Munro et al. [1] and a set of 84 reactions from RetroTransformDB [2] represented as SMIRKS linear notations (see file <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/SMIRKS-RetroDB.txt?versionId=2fa6cb02-013a-4f69-b89a-c18d03b09222">SMIRKS-RetroDB.txt</a>). In both software tools (RDKit and Ambit-SMIRKS), each reaction was applied for all compounds at all possible sites thus performing more than 46000 SMIRKS transformations. For the purpose of comparison, Ambit-SMIRKS was applied in mode ALL with a single copy of the products for each reaction site. The java code for Ambit-SMIRKS test is available in file <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/TestAmbitSmirks.java?versionId=fcb01365-b869-4801-9c5d-43cf8765103a">TestAmbitSmirks.java</a> and respectively python code for RDKit test is present in <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/rdkit-smirks-test-02.py?versionId=d42ab1c5-4c1f-4fad-8457-a6e8a8f4a3ec">rdkit-smirks-test-02.py</a>. In order to run the tests, Ambit dependency modules (version 3.2.0) are required (see more about Ambit at https://ambit.sf.net/) as well as RDKit (release 2018.03) installation is needed (see http://www.rdkit.org/).</p> <p>The tests were performed on a PC computer (Intel/Core i5-8250U, 1.6GHz/12 GB RAM), under Win10 Operating system. The calculations took about 30 seconds for RDKit software and about 40 seconds for Ambit-SMIRKS. &nbsp;The computational time for both software includes the SMIRKS parsing and reaction application as well as molecule preprocessing and file operations. Each algorithm was run 3 times. Detail timing info is present in file <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/time-stat.txt?versionId=57e06cd1-d454-4bc0-a369-13b1e4d6150f">time-stat.txt</a>.</p> <p>The raw data outputs for both software tools respectively are stored in files: <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/rdkit-out.txt?versionId=2fa730f9-6ed8-4e51-a42c-c47799368f3b">rdkit-out.txt </a>and&nbsp; <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/ambit-out-no-eq-filter.txt?versionId=0519c609-06ae-4da8-b3c8-28a316d517d0">ambit-out-no-eq-filter.txt </a></p> <p>The generated output files are constructed from blocks for each SMIRKS in the following format:</p> <p>##smirks-number &lt;SMIRKS&gt;</p> <p>&lt;smiles1&gt; --&gt; &lt;number of reaction sites&gt; &lt;products1&gt;, &lt;products2&gt;, &hellip;</p> <p>&lt;smiles2&gt; --&gt; &lt;number of reaction sites&gt; &lt;products1&gt;, &lt;products2&gt;, &hellip;</p> <p>&hellip;</p> <p>&lt;smiles545&gt; --&gt; &lt;number of reaction sites&gt; &lt;products1&gt;, &lt;products2&gt;, &hellip;</p> <p>On the base of generated raw test data, comparison statistics was summarized in file&nbsp; <a href="https://zenodo.org/api/files/b8c44e94-55a6-48b9-bd11-ff5cee71e55a/compare-ambit-rdkit.xlsx?versionId=5d4d6673-a9f0-47fd-8512-86112b3483f6">compare-ambit-rdkit.xlsx </a>containing following columns: <strong>SMILES</strong> &ndash; target molecule smiles, <strong>smirks_num</strong> &ndash; the index of reaction SMIRKS applied against the target, <strong>Ambit-NEF</strong> &ndash; number of reacted sites in the target molecule for Ambit algorithm, <strong>RDKit</strong> - number of reacted sites in the target molecule for RDKit algorithm , <strong>Diff</strong> &ndash; absolute difference the number reacted sites in Ambit and RDKit, <strong>FlagDiff</strong> &ndash; it is 1 (true) if the <strong>Diff</strong> is non zero, <strong>FlagRDKitReact</strong> &ndash; it is 1 (true) if at least one site is reacted in the target molecule by RDKit tool (i.e. RDKit column values &gt; 0), <strong>FlagAmbitReact</strong> - it is 1 (true) if at least one site is reacted in the target molecule by Ambit-SMIRKS tool (i.e. Ambit-NEF column values &gt; 0).</p> <p>Out of 46410 tests, 6096 test reactions were successfully applied for at least one site in Ambit-SMIRKS (i.e. the value in column Ambit-NEF is not zero) and 5729 reactions were successfully applied for at least one site in RDKit accordingly (i.e. the value in column RDKit is not zero). The obtained total number of reacted sites for Ambit-SMIRKS and RDKit is 41453 and 40782 respectively. We have performed statistics of the number of reacted sites for both software tools and differences were observed for 436 reactions. From our analysis we may infer that the observed differences are mainly due to different treatment of equivalent molecules sites and some small differences of the internal presentation of the molecules and the chemical reactions on both software packages.&nbsp;</p> <p>[1] Munro I., Ford RA, Kennepohl E, Sprenger J. Correlation of structural class with no-observed-effect-levels: a proposal for establishing a threshold of concern. Food Chem Toxicol. 1996;34:829&ndash;867.</p> <p>[2] https://doi.org/10.5281/zenodo.1209313</p>

opencc-by-4.0Jul 2018View details →
zenodo44/100

Italian Lexical Simplification Benchmark

<p>The corpus is a manually created benchmark to evaluate the performance of Italian lexical simplification systems. It contains 901 pairs of complex sentences and their simplified version&nbsp;at the lexical level (i.e. replacement of a difficult term or phrase with a simpler synonym). The dataset and a system using the benchmark are&nbsp;described in the paper &quot;The impact of phrases on Italian lexical simplification&quot;&nbsp;&nbsp;<a href="https://zenodo.org/record/1048874">https://zenodo.org/record/1048874</a></p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

Supporting Jupyter Python notebook for "A new class of efficient randomized benchmarking protocols"

<p>Python notebook containing the code used to generate the data for figure 2&nbsp;in the appendix of &quot;A new class of efficient randomized benchmarking protocols&quot; (arXiv:1806.02048).</p>

opencc-by-4.0Jan 2019View details →
zenodo44/100

Simulated Arabidopsis thaliana sequencing datasets for chloroplast assembler benchmarking

<p><strong>Changes</strong></p> <ul> <li>Fixed non-circular sampling from chloroplast and mitochondrion in version 1.1.0</li> <li>Fixed off-by-one error in reverse read in version 1.0.0</li> </ul> <p><strong>Purpose and Documentation</strong></p> <p>See: <a href="https://github.com/chloroExtractorTeam/benchmark">github.com/chloroExtractorTeam/benchmark</a></p> <p><strong>Original data</strong><br> The original <em>Arabidopsis thaliana </em>sequences were downloaded from TAIR:&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;</p> <p>The Arabidopsis Information Resource (<a>TAIR</a>) on www.arabidopsis.org, Mar 22, 2019 available under the <a href="http://www.arabidopsis.org/doc/about/tair_terms_of_use/417">TAIR Terms of Use</a>&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp;&nbsp; &nbsp;</p> <p><em>Tanya Z. Berardini, Leonore Reiser, Donghui Li, Yarik Mezheritsky, Robert Muller, Emily Strait and Eva Huala. &quot;The Arabidopsis Information Resource: Making and mining the &quot;gold standard&quot; annotated reference plant genome.&quot;&nbsp;&nbsp;&nbsp; genesis 2015 <a href="https://doi.org/10.1002/dvg.22877">doi:10.1002/dvg.22877</a></em></p> <p><strong>Programs used to generate this data</strong><br> &nbsp;- <a href="https://github.com/shenwei356/seqkit">seqkit</a> (v0.10.1): Shen W, Le S, Li Y, Hu F (2016) &quot;SeqKit: A Cross-Platform and Ultrafast Toolkit for FASTA/Q File Manipulation.&quot; PLOS ONE 11(10): e0163962. <a href="https://doi.org/10.1371/journal.pone.0163962">doi:10.1371/journal.pone.0163962</a></p> <p>&nbsp;</p>

opencc-by-4.0Apr 2019View details →
zenodo44/100

List.MID: A MIDI-Based Benchmark for Evaluating RDF Lists

<p>Linked lists represent a countable number of ordered values, and are among the most important abstract data types in computer science. With the advent of RDF as a highly expressive knowledge representation language for the Web, various implementations for RDF lists have been proposed. Yet, there is no benchmark so far dedicated to evaluate the performance of triple stores and SPARQL query engines on dealing with ordered linked data. Moreover, essential tasks for evaluating RDF lists, like generating datasets containing RDF lists of various sizes, or generating the same RDF list using different modelling choices, are cumbersome and unprincipled. In this paper, we propose List.MID, a systematic benchmark for evaluating systems serving RDF lists. List.MID consists of a dataset generator, which creates RDF list data in various models and of different sizes; and a set of SPARQL queries. The RDF list data is coherently generated from a large, community-curated base collection of Web MIDI files, rich in lists of musical events of arbitrary length. We describe the List.MID benchmark, and discuss its impact and adoption, reusability, design, and availability.</p>

opencc-by-sa-4.0Dec 2018View details →
zenodo44/100

Benchmark protocol for exoplanet forward model and retrieval

<p>Benchmark protocol for giant exoplanet atmosphere tools, presented in Baudino et al. 2017 <a href="https://doi.org/10.3847/1538-4357/aa95be">https://doi.org/10.3847/1538-4357/aa95be</a></p> <p>The original data to reproduice the protocol are used in a jupyter notebook &quot;Tutorial.ipynb&quot; including all the plot routines to help to compare with you own models</p>

opencc-by-4.0Nov 2017View details →
zenodo44/100

Improved upper bounds for permutation flowshop scheduling benchmarks (Taillard and VRF)

<p>Optimal makespans and permutation schedules (found and proven optimal by Branch-and-Bound) for Taillard instances Ta112, Ta116 (500 jobs, 20 machines) and 74 instances of the VRF benchmark.</p>

opencc-by-4.0Nov 2019View details →
zenodo44/100

A SAT Benchmark Suite for LTL Specification Sketching

<h1>LTL_Sketcher-SAT_Benchmark</h1> <p>This repository contains a set of formulas in Propositional Boolean Logic.<br>These formulas are generated during the execution of our <a href="https://github.com/rajarshi008/LTLSketcher/tree/master" target="_blank" rel="noopener">LTLSketcher tool</a>.<br>Given an LTL sketch (i.e., a partial LTL formula) and a sample (i.e., a set of program executions labeled desired and undesired), the tool solves the LTL sketching problem, i.e., complete the sketch to a specification consistent with the data. (feel freet to check out our <a href="https://link.springer.com/chapter/10.1007/978-3-031-45332-8_2" target="_blank" rel="noopener">paper</a> for more information on this problem)<br>In essence, this is done by reducing the problem to a series of formulas in Propositional Boolean Logic and checking their satisfiability.</p> <h2>Naming convention:</h2> <p>This repository contains each formula both in the DIMACS and SMTLib format.<br>Each file follows the same naming convention:</p> <p><em>type__sample-file__sketch__size__algorithm-configuration__satisifability</em></p> <p><em>type</em>: indicates whether the formula is stored in the DIMACS or SMTLib format<br><em>sample-file</em>: refers to the sample (cf., <a href="https://github.com/rajarshi008/LTLSketcher/tree/master/experiment_results/generated_files/final_benchmark" target="_blank" rel="noopener">here</a>) used by the LTLSketcher tool<br><em>sketch</em>: refers to the sketch (cf., See experimental evaluation of our <a href="https://link.springer.com/chapter/10.1007/978-3-031-45332-8_2" target="_blank" rel="noopener">paper</a>) used by the LTLSketcher tool<br><em>size</em>: refers to the size of the complete solution (i.e., the number of subformulas of the complete specification)<br><em>algorithm-configuration</em>: our algorithm can be extended by two heuristics (BMC and suffix), this indicates which combination of heuristics was used (none, either one of the two, both)<br><em>satisfiability</em>: indicates whether the formula is satisfiable or not</p>

opencc-by-4.0Apr 2024View details →
zenodo44/100

UnientrezDB: Large-scale Gene Ontology Annotation Dataset and Evaluation Benchmarks with Unified Entrez Gene Identifiers

<p>Our work focuses on providing a comprehensive dataset and benchmarks for evaluating gene ontology annotations using a unified system of Entrez Gene Identifiers.</p>

opencc-by-4.0Aug 2024View details →
zenodo44/100

Datasets used in the benchmarking study of MR methods

<p>We conducted a benchmarking analysis of 16 summary-level data-based MR methods for causal inference with five real-world genetic datasets, focusing on three key aspects: type I error control, the accuracy of causal effect estimates, replicability, and power.</p> <p>The datasets used in the MR benchmarking study can be downloaded here:</p> <ol> <li>"dataset-GWASATLAS-negativecontrol.zip":&nbsp; the GWASATLAS dataset for evaluation of type I error control in confounding scenario (a): Population stratification</li> <li>"dataset-NealeLab-negativecontrol.zip": the Neale Lab dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-PanUKBB-negativecontrol.zip": the Pan UKBB dataset for evaluation of type I error control in confounding scenario (a): Population stratification;</li> <li>"dataset-Pleiotropy-negativecontrol": the dataset&nbsp; used for evaluation of type I error control in confounding scenario (b): Pleiotropy;</li> <li>"dataset-familylevelconf-negativecontrol.zip": the dataset used for evaluation of type I error control in confounding scenario (c): Family-level confounders;</li> <li>"dataset_ukb-ukb.zip": the dataset used for evaluation of the accuracy of causal effect estimates;</li> <li>"dataset-LDL-CAD_clumped.zip": the dataset used for evaluation of replicability and power;</li> </ol> <p>Each of the datasets contains the following files:</p> <ol> <li>&nbsp;"Tested Trait pairs": the exposure-outcome trait pairs to be analyzed;</li> <li>"MRdat" refers to the summary statistics after performing IV selection (p-value &lt; 5e-05) and PLINK LD clumping with a clumping window size of 1000kb and an r^2 threshold of 0.001.</li> <li>"bg_paras" are the estimated background parameters "Omega" and "C" which will be used for MR estimation in MR-APSS.</li> </ol> <p>Note:</p> <ol> <li>The formatted dataset after quality control can be accessible at our GitHub website (https://github.com/YangLabHKUST/MRbenchmarking).</li> <li>The details on quality control of GWAS summary statistics, formatting GWASs, and LD clumping for IV selection can be found on the MR-APSS software tutorial on the MR-APSS&nbsp;&nbsp;website (https://github.com/YangLabHKUST/MR-APSS).</li> <li>R code for running MR methods is also available at https://github.com/YangLabHKUST/MRbenchmarking.</li> </ol>

opencc-by-4.0Jan 2024View details →

ScienceDex guides

Understand access before you commit

These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.

Compare curated datasets

Allen Brain Atlas

Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.

allen-brain-atlas
neuroscienceopenDocumentation, web resources, and API references are available online.
Last verified 2026-04-30Open record

Annotated Behaviour and Observability Dataset (ABODe)

ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.

abode-home-cage
behavioral-neuroscienceopenThe DataShare record exposes download links for annotations, documentation, license text, and the zipped per-snippet data directory.
Last verified 2026-04-30Open record

DANDI Archive for NWB datasets

DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.

dandi-nwb
electrophysiologyopenPublished Dandiset metadata and archive endpoints are available through the production DANDI API.
Last verified 2026-04-30Open record

International Brain Laboratory public data

The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.

ibl
behavioral-neuroscienceopenPublic sessions can be searched and loaded from the IBL public data server through ONE.
Last verified 2026-04-29Open record

OpenNeuro

OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.

openneuro
neuroscienceopenPublished datasets are available on demand over the internet.
Last verified 2026-04-29Open record