Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
237
datasets available to search
ShareScore release 0.9.0
Dataset results
237 results for “OCR”
ICDAR2015 competition on smartphone document capture and OCR (SmartDoc) - Challenge 2
<p><strong>ICDAR2015 competition on smartphone document capture and OCR (SmartDoc)</strong></p> <p><strong>Challenge 2: MOBILE OCR COMPETITION</strong></p> <p>The goal of the competition is to extract the textual content from document images which are captured by mobile phones. The images are taken under varying conditions to provide a challenging input. The dataset was prepared for ICDAR2015-SmartDoc competition. For more details about the dataset please visit the competition's website:</p> <p>https://sites.google.com/site/icdar15smartdoc/home</p> <p>http://smartdoc.univ-lr.fr</p> <p>You may also refer to the following paper for more details on the ICDAR2015-SmartDoc competition:</p> <p>Jean-Christophe Burie, Joseph Chazalon, Mickaël Coustaty, Sébastien Eskenazi, Muhammad Muzzamil Luqman, Maroua Mehri, Nibal Nayef, Jean-Marc OGIER, Sophea Prum and Marçal Rusinol: “ICDAR2015 Competition on Smartphone Document Capture and OCR (SmartDoc)”, In 13th International Conference on Document Analysis and Recognition (ICDAR), 2015.</p> <p><strong>If you use this dataset, please send us a short email at <icdar.smartdoc (at) gmail.com> to tell us why it was useful to you, and whether you have results or publications we can reference on our website. Thank you!</strong></p>
NewsEye / READ OCR training dataset from Austrian Newspapers (19th C.)
<p>The dataset comprises Austrian newspaper pages from 19th and early 20th century with carefully corrected text. The page images were provided by the <a href="http://onb.ac.at/">Austrian National Library</a> and comprise 148 pages (training set) and 13 pages (validation set). The data are formed according to the PAGE format (cf. Cf. <a href="https://github.com/PRImA-Research-Lab/PAGE-XML/">https://github.com/PRImA-Research-Lab/PAGE-XML/</a>) and were produced with the <a href="http://read.transkribus.eu/">Transkribus </a>platform with support of the <a href="http://newseye.eu/">NewsEye</a> and the <a href="http://read.transkribus.eu/">READ </a>project.</p>
Dataset of ICDAR 2019 Competition on Post-OCR Text Correction
<p><strong>Corpus for the ICDAR2019 Competition on Post-OCR Text Correction (October 2019)</strong><br> Christophe Rigaud, Antoine Doucet, Mickael Coustaty, Jean-Philippe Moreux<br> <a href="http://l3i.univ-larochelle.fr/ICDAR2019PostOCR">http://l3i.univ-larochelle.fr/ICDAR2019PostOCR</a><br> -------------------------------------------------------------------------------</p> <p>These are the supplementary materials for the ICDAR 2019 paper <em><a href="https://zenodo.org/record/3459116">ICDAR 2019 Competition on Post-OCR Text Correction</a></em></p> <p>Please use the following citation:</p> <pre><code>@inproceedings{rigaud2019pocr,</code> <code> </code><code>title="ICDAR 2019 Competition on Post-OCR Text Correction",</code> <code> </code><code>author={Rigaud, Christophe and Doucet, Antoine and Coustaty, Mickael and Moreux, Jean-Philippe},</code> <code> </code><code>year={2019},</code> <code> </code><code>booktitle={Proceedings of the 15th International Conference on Document Analysis and Recognition (2019)}</code> <code> </code><code>}</code></pre> <p> </p> <p><strong>Description</strong><br> The corpus accounts for 22M OCRed characters along with the corresponding Gold Standard (GS). The documents come from different digital collections available, among others, at the National Library of France (BnF) and the British Library (BL). The corresponding GS comes both from BnF's internal projects and external initiatives such as Europeana Newspapers, IMPACT, Project Gutenberg, Perseus and Wikisource.</p> <p><strong>Repartition of the dataset</strong><br> - <em>ICDAR2019_Post_OCR_correction_training_18M.zip</em>: 80% of the full dataset, provided to train participants' methods.<br> - <em>ICDAR2019_Post_OCR_correction_evaluation_4M</em>: 20% of the full dataset used for the evaluation (with Gold Standard made publicly after the competition).<br> - <em>ICDAR2019_Post_OCR_correction_full_22M</em>: full dataset made publicly available after the competition.</p> <p><strong>Special case for Finnish language</strong><br> Material from the National Library of Finland (<em>Finnish dataset FI > FI1</em>) are not allowed to be re-shared on other website. Please follow these guidelines to get and format the data from the original website.</p> <p>1. Go to <a href="https://digi.kansalliskirjasto.fi/opendata/submit?set_language=en">https://digi.kansalliskirjasto.fi/opendata/submit?set_language=en</a>;<br> 2. Download <em>OCR Ground Truth Pages (Finnish Fraktur) [v1](4.8GB)</em> from <em>Digitalia (2015-17)</em> package;<br> 3. Convert the Excel file "<em>~/metadata/nlf_ocr_gt_tescomb5_2017.xlsx</em>" as Comma Separated Format (.csv) by using <em>save as</em> function in a spreadsheet software (e.g. Excel, Calc) and copy it into "<em>FI/FI1/HOWTO_get_data/input/</em>";<br> 4. Go to "<em>FI/FI1/HOWTO_get_data/</em>" and run "<em>script_1.py</em>" to generate the <em>full</em> "<em>FI1</em>" dataset in "<em>output/full/</em>";<br> 4. Run "<em>script_2.py</em>" to split the "<em>output/full/</em>" dataset into "<em>output/training/</em>" and "<em>output/evaluation/</em>" sub sets.<br> At the end of the process, you should have a "<em>training</em>", "<em>evaluation</em>" and "<em>full</em>" folder with 1579528, 380817 and 1960345 characters respectively.</p> <p><br> <strong>Licenses: free to use for non-commercial uses, according to sources in details</strong><br> - BG1: IMPACT - National Library of Bulgaria: CC BY NC ND<br> - CZ1: IMPACT - National Library of the Czech Republic: CC BY NC SA<br> - DE1: Front pages of Swiss newspaper NZZ: Creative Commons Attribution 4.0 International (<a href="https://zenodo.org/record/3333627">https://zenodo.org/record/3333627</a>)<br> - DE2: IMPACT - German National Library: CC BY NC ND<br> - DE3: GT4Hist-dta19 dataset: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE4: GT4Hist - EarlyModernLatin: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE5: GT4Hist - Kallimachos: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE6: GT4Hist - RefCorpus-ENHG-Incunabula: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - DE7: GT4Hist - RIDGES-Fraktur: CC-BY-SA 4.0 (<a href="https://zenodo.org/record/1344132">https://zenodo.org/record/1344132</a>)<br> - EN1: IMPACT - British Library: CC BY NC SA 3.0<br> - ES1: IMPACT - National Library of Spain: CC BY NC SA<br> - FI1: National Library of Finland: no re-sharing allowed, follow the above section to get the data. (<a href="https://digi.kansalliskirjasto.fi/opendata">https://digi.kansalliskirjasto.fi/opendata</a>)<br> - FR1: HIMANIS Project: CC0 (<a href="https://www.himanis.org/">https://www.himanis.org</a>)<br> - FR2: IMPACT - National Library of France: CC BY NC SA 3.0<br> - FR3: RECEIPT dataset: CC0 (<a href="http://findit.univ-lr.fr/">http://findit.univ-lr.fr</a>)<br> - NL1: IMPACT - National library of the Netherlands: CC BY<br> - PL1: IMPACT - National Library of Poland: CC BY<br> - SL1: IMPACT - Slovak National Library: CC BY NC</p> <p>Text post-processing such as cleaning and alignment have been applied on the resources mentioned above, so that the Gold Standard and the OCRs provided are not necessarily identical to the originals.</p> <p><br> <strong>Structure</strong><br> - **Content** [<em>./lang_type/sub_folder/#.txt</em>]<br> - "<em>[OCR_toInput] </em>" => Raw OCRed text to be de-noised.<br> - "<em>[OCR_aligned] </em>" => Aligned OCRed text.<br> - "<em>[ GS_aligned] </em>" => Aligned Gold Standard text.</p> <p>The aligned OCRed/GS texts are provided for training and test purposes. The alignment was made at the character level using "<em>@</em>" symbols. "<em>#</em>" symbols correspond to the absence of GS either related to alignment uncertainties or related to unreadable characters in the source document. For a better view of the alignment, make sure to disable the "word wrap" option in your text editor.</p> <p>The Error Rate and the quality of the alignment vary according to the nature and the state of degradation of the source documents. Periodicals (mostly historical newspapers) for example, due to their complex layout and their original fonts have been reported to be especially challenging. In addition, it should be mentioned that the quality of Gold Standard also varies as the dataset aggregates resources from different projects that have their own annotation procedure, and obviously contains some errors.</p> <p><br> <strong>ICDAR2019 competition</strong><br> Information related to the tasks, formats and the evaluation metrics are details on :<br> <a href="https://sites.google.com/view/icdar2019-postcorrectionocr/evaluation">https://sites.google.com/view/icdar2019-postcorrectionocr/evaluation</a></p> <p><br> <strong>References</strong><br> - IMPACT, European Commission's 7th Framework Program, grant agreement 215064<br> - Uwe Springmann, Christian Reul, Stefanie Dipper, Johannes Baiter (2018). Ground Truth for training OCR engines on historical documents in German Fraktur and Early Modern Latin.<br> - <a href="https://digi.nationallibrary.fi/">https://digi.nationallibrary.fi</a> , Wiipuri, 31.12.1904, Digital Collections of National Library of Finland<br> - EU Horizon 2020 research and innovation programme grant agreement No 770299</p> <p><br> <strong>Contact</strong><br> - christophe.rigaud(at)univ-lr.fr<br> - antoine.doucet(at)univ-lr.fr<br> - mickael.coustaty(at)univ-lr.fr<br> - jean-philippe.moreux(at)bnf.fr</p> <p>L3i - University of la Rochelle, <a href="http://l3i.univ-larochelle.fr/">http://l3i.univ-larochelle.fr</a><br> BnF - French National Library, <a href="http://www.bnf.fr/">http://www.bnf.fr</a></p>
Noisy OCR Dataset (NOD)
<p>This dataset contains 18,504 images of English and Arabic documents with ground truth for use in OCR benchmarking. It consists of two collections, "Old Books" (English) and "Yarmouk" (Arabic), each of which contains an image set reproduced in 44 versions with different types and degrees of artificially generated noise. The dataset was originally developed for Hegghammer (2021).</p> <p><strong>Source images</strong></p> <p>The seed of the English collection was the "Old Books Dataset" (Barcha 2017), a set of 322 page scans from English-language books printed between 1853 and 1920. The seed of the Arabic collection was a randomly selected subset of 100 pages from the "Yarmouk Arabic OCR Dataset" (Abu Doush et al. 2018), which consists of 4,587 Arabic Wikipedia articles printed to paper and scanned to PDF.</p> <p><strong>Artificial noise application</strong></p> <p>The dataset was created as follows:<br> - First a greyscale version of each image was created, so that there were two versions (colour and greyscale) with no added noise. <br> - Then six ideal types of image noise --- "blur", "weak ink", "salt and pepper", "watermark", "scribbles", and "ink stains" --- were applied both to the colour version and the binary version of the images, thus creating 12 additional versions of each image. The R code used to generate the noise is included in the repository.<br> - Lastly, all available combinations of *two* noise filters were applied to the colour and binary images, for an additional 30 versions. </p> <p>This yielded a total of 44 image versions divided into three categories of noise intensity: 2 versions with no added noise, 12 versions with one layer of noise, and 30 versions with two layers of noise. This amounted to an English corpus of 14,168 documents and an Arabic corpus of 4,400 documents. </p> <p>The compressed archive is ~26 GiB, and the uncompressed version is ~193 GiB. See <a href="http://www.e7z.org/open-lzma-tlz.htm">this link</a> for how to unzip <em>.tar.lzma</em> files. </p> <p><strong>References:</strong></p> <p>Barcha, Pedro. 2017. “Old Books Dataset.” GitHub Repository. <em>GitHub</em>. https:<br> //github.com/PedroBarcha/old-books-dataset.</p> <p>Doush, Iyad Abu, Faisal AlKhateeb, and Anwaar Hamdi Gharibeh. 2018. “Yarmouk<br> Arabic OCR Dataset.” In <em>2018 8th International Conference on Computer Science<br> and Information Technology (CSIT)</em>, 150–54. IEEE.</p> <p>Hegghammer, Thomas. 2021. "OCR with Tesseract, Amazon Textract, and Google Document AI: A Benchmarking Experiment". <em>Socarxiv</em>. https://osf.io/preprints/socarxiv/6zfvs</p>
Patrologia Graeca (OCR ground truth)
<p>Ground truth manually produced within the scope of the CGPG project (Calfa GREgORI Patrologia Graeca), led by Jean-Marie Auwers (UCLouvain), that aim to OCRize the remaining non-digital versions of the Patrologia Graeca volumes. This dataset compiles annotations from 2021 to 2022, and has been used for the Programming Historian online lesson "Transcription automatisée de graphies non latines" (link to come).</p> <p>The dataset contains a set of 100 images from the Patrologia Graeca, with their corresponding pageXML files. Annotation has been performed with the <a href="https://vision.calfa.fr">Calfa Vision platform</a>.</p> <p>Different level of annotations are proposed to overcome two tasks: (task1) detection of text regions in Greek (annotations at the region level only) and (task 2) recognition of ancient polytonic Greek (annotations at the line level).</p> <p><strong>Task 1:</strong><br> col_greek: 52<br> col_lat: 54<br> footnotes: 27<br> titles: 9</p> <p><strong>Task 2:</strong><br> lines: 2.579</p> <p>Final model of layout analysis is freely usable on the <a href="https://vision.calfa.fr">Calfa Vision platform</a>, by choosing the "Greek printed (Patrologia Graeca)" type of project.</p> <p>The project is sponsored by the ASBL <em>Byzantion</em>, the Fondation <em>Sedes Sapientiae</em>, the Institut <em>Religions, Spiritualités, Cultures, Sociétés</em> (RSCS, UCLouvain) and the <em>Centre d'études orientales</em> (CIOL, UCLouvain) and by a generous donor who wishes to remain anonymous. Other sponsors have recently expressed their willingness to support the project.</p>
HHD-Ethiopic: A Historical Handwritten Dataset for Ethiopic OCR
<p>HHD-Ethiopic is a historical handwritten dataset for Ethiopic text-image recognition.</p>
Historical German Children's Playbooks - 6 Digitized Books with Images, OCR-Fulltext, and Named Entity Recognition
<p>The dataset consists of 6 digitized books with 1750 images and OCR-fulltext.</p> <p>Additionally, named entity recognition has been carried out on basis of flair's de-ner model, see https://github.com/flairNLP for details.</p>
DATASET PARA ANÁLISE UNIVARIADA DE OCR
<p>Repositório de dados contendo 52.525 imagens sintéticas para avaliação univariada de sistemas de Reconhecimento Óptico de Caracteres e estudos sobre detecção de fake news.</p>
Tesseract OCR models for the Alsatian dialects
<p>This dataset provides trained Tesseract (<a href="https://github.com/tesseract-ocr/tesseract">https://github.com/tesseract-ocr/tesseract</a>) OCR models for the Alsatian dialects. These models were developed in the context of the RESTAURE project, funded by the French ANR. </p> <p>Two models are provided :</p> <p>The first model, ISKO_2015, has been presented in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01252241">https://hal.archives-ouvertes.fr/hal-01252241</a>. The Tesseract model has been trained using the jTessBoxEditor tool (<a href="http://vietocr.sourceforge.net/training.html">http://vietocr.sourceforge.net/training.html</a>), Version 1.4 (2 May 2015), based on images automatically generated from the training texts (excerpts from 7 different printed works, totalling about 9,000 words). The generation of the images used a 36pt font size, and two fonts were used (Arial and Times New Roman), with their normal and italic variants.<br> The Tesseract model (gsw.traineddata) can be used with Tesseract 3.0x.</p> <p>The second model, 2018, has been trained for Tesseract 4.0x, using jTessBoxEditor version 2.0.1 (28 July 2018). Again, images were automatically generated from the training text. The training text is different from the one used for the ISKO_2015 model and is "artificial", in the sense that it has been built by appending word n-grams extracted from a large variety of published texts in Alsatian, for a time period spanning 2 centuries and for different text genres. The images corresponding to this training text have been automatically generated with the Tesseract text2image tool, using the following parameters: --ptsize=36 --leading=20. The fonts used are listed in the gsw.font_properties file.</p> <p>Dictionary data has also been used for training. We conflated Alsatian words found in several lexicons and corpora:</p> <ul> <li>Lexicons produced by the OLCA (Office pour la Langue et les Cultures d'Alsace et de Moselle): <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Lexicon from a Wiktionary user page: <a href="http://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais">https://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais</a></li> <li>Lexicon from the ACPA association: <a href="http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm">http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm</a></li> <li>Chronicles published by Raymond Matzen in the local newspaper "Les Dernières Nouvelles d'Alsace"</li> <li>Transcriptions of television shows found in Erhart, P. (2012). <em>Les dialectes dans les médias: quelle image de l’Alsace véhiculent-ils dans les émissions de la télévision régionale?</em>, Université de Strasbourg, <a href="http://www.theses.fr/167563386">http://www.theses.fr/167563386</a></li> <li>French-Alsatian parallel corpus provided by the OLCA</li> <li>Excerpts from Adolf, P. (2006). <em>Dictionnaire comparatif multilingue: français-allemand-alsacien-anglais.</em>, Strasbourg, France, Midgard, 2006, 373 p.</li> </ul> <p>The Tesseract models can be used for instance using the gImageReader tool (<a href="https://github.com/manisandro/gImageReader">https://github.com/manisandro/gImageReader</a>), which provides a graphical user interface for the Tesseract tool. </p> <p>When evaluated against the same test corpus (prose by Marie Hart, theater and poetry by Gustave Stokopf and prose by Charles Zumstein, totalling about 4,900 words), both models achieve roughly the same performance levels. Usually, even better performance levels can be achieved by combining the Alsatian-specific model with the French and German models available for Tesseract (available from <a href="https://github.com/tesseract-ocr/tessdata">https://github.com/tesseract-ocr/tessdata</a>)</p>
DBNL OCR Data set
<p>A set of 220 books digitised by the Dutch DBNL (<a href="https://dbnl.org/">https://dbnl.org/</a>). The set contains the original OCR output in .txt and the corrected version in TEI.</p>
Estonian Historical Newspaper Crowdsourced OCR Corrections
<div> <div>This dataset consists of newspaper articles from the National Library of Estonia's DIGAR archive and their respective crowdsourced corrections.</div> </div>
Molecule OCR Real images Dataset
<p>Test dataset from paper <strong>Image2SMILES: Transformer-based Molecular Optical Recognition Engine</strong>. The dataset contains 296 structures: images and Functional Groups SMILES (FG-SMILES). The structures were extracted from 24 papers, which were selected from each volume of Journal of Organic Chemistry (2020). </p>
FX2173 ocr-4(tm2173)IV | 2010-03-19T10:34:04+00:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=eu5_j1uLO8M</li> <li><b>strain</b> : FX2173</li> <li><b>timestamp</b> : 2010-03-19T10:34:04+00:00</li> <li><b>gene</b> : ocr-4</li> <li><b>chromosome</b> : IV</li> <li><b>allele</b> : tm2173</li> <li><b>strain_description</b> : ocr-4(tm2173)IV</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : clockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4 (tm2173) on food L_2010_03_19__10_34_04___1___6</li> <li><b>total time (s)</b> : 899.612</li> <li><b>frames per second</b> : 25.3165</li> <li><b>video micrometers per pixel</b> : 4.29558</li> <li><b>number of segmented skeletons</b> : 19558</li> </ul>
LX982 ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V | 2010-03-26T12:35:50+00:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=O3KTpS_LFy4</li> <li><b>strain</b> : LX982</li> <li><b>timestamp</b> : 2010-03-26T12:35:50+00:00</li> <li><b>gene</b> : ocr-1;ocr-2;ocr-4</li> <li><b>chromosome</b> : IV;V</li> <li><b>allele</b> : ok132;ak47;vs137</li> <li><b>strain_description</b> : ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : clockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4 (vs137); ocr-2 (9447); ocr-1 (ok134) on food L_2010_03_26__12_35_50___1___8</li> <li><b>total time (s)</b> : 898.667</li> <li><b>frames per second</b> : 25.974</li> <li><b>video micrometers per pixel</b> : 4.29558</li> <li><b>number of segmented skeletons</b> : 19592</li> </ul>
LX982 ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V | 2010-03-25T15:22:20+00:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=7PDVRAoCZX4</li> <li><b>strain</b> : LX982</li> <li><b>timestamp</b> : 2010-03-25T15:22:20+00:00</li> <li><b>gene</b> : ocr-1;ocr-2;ocr-4</li> <li><b>chromosome</b> : IV;V</li> <li><b>allele</b> : ok132;ak47;vs137</li> <li><b>strain_description</b> : ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : clockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4 (vs137); ocr-2 (9447); ocr-1 (ok134) on food L_2010_03_25__15_22_20___1___8</li> <li><b>total time (s)</b> : 898.668</li> <li><b>frames per second</b> : 25.7069</li> <li><b>video micrometers per pixel</b> : 4.29558</li> <li><b>number of segmented skeletons</b> : 19374</li> </ul>
FX2173 ocr-4(tm2173)IV | 2010-03-26T14:52:17+00:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=atrqCRO4KYs</li> <li><b>strain</b> : FX2173</li> <li><b>timestamp</b> : 2010-03-26T14:52:17+00:00</li> <li><b>gene</b> : ocr-4</li> <li><b>chromosome</b> : IV</li> <li><b>allele</b> : tm2173</li> <li><b>strain_description</b> : ocr-4(tm2173)IV</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : clockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4 (tm2173) on food L_2010_03_26__14_52_17___1___10</li> <li><b>total time (s)</b> : 898.131</li> <li><b>frames per second</b> : 25.641</li> <li><b>video micrometers per pixel</b> : 4.29558</li> <li><b>number of segmented skeletons</b> : 19512</li> </ul>
LX982 ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V | 2010-07-06T12:26:06+01:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=DZm2qgAV-PQ</li> <li><b>strain</b> : LX982</li> <li><b>timestamp</b> : 2010-07-06T12:26:06+01:00</li> <li><b>gene</b> : ocr-1;ocr-2;ocr-4</li> <li><b>chromosome</b> : IV;V</li> <li><b>allele</b> : ok132;ak47;vs137</li> <li><b>strain_description</b> : ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : clockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4 (vs137); ocr-2 (9447); ocr-1 (ok134) on food L_2010_07_06__12_26_06___1___8</li> <li><b>total time (s)</b> : 898.399</li> <li><b>frames per second</b> : 25.3807</li> <li><b>video micrometers per pixel</b> : 4.36527</li> <li><b>number of segmented skeletons</b> : 18914</li> </ul>
LX982 ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V | 2010-07-08T10:43:15+01:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=mfzaJ4B07bg</li> <li><b>strain</b> : LX982</li> <li><b>timestamp</b> : 2010-07-08T10:43:15+01:00</li> <li><b>gene</b> : ocr-1;ocr-2;ocr-4</li> <li><b>chromosome</b> : IV;V</li> <li><b>allele</b> : ok132;ak47;vs137</li> <li><b>strain_description</b> : ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : anticlockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4 (vs137); ocr-2 (9447); ocr-1 (ok134) on food R_2010_07_08__10_43_15___1___2</li> <li><b>total time (s)</b> : 899.328</li> <li><b>frames per second</b> : 26.0417</li> <li><b>video micrometers per pixel</b> : 4.36527</li> <li><b>number of segmented skeletons</b> : 18714</li> </ul>
LX950 ocr-4(vs137)IV | 2010-04-23T15:19:23+01:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=xMBKHt97HLo</li> <li><b>strain</b> : LX950</li> <li><b>timestamp</b> : 2010-04-23T15:19:23+01:00</li> <li><b>gene</b> : ocr-4</li> <li><b>chromosome</b> : IV</li> <li><b>allele</b> : vs137</li> <li><b>strain_description</b> : ocr-4(vs137)IV</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : anticlockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4 (vs137) on food R_2010_04_23__15_19_23___8___12</li> <li><b>total time (s)</b> : 899.73</li> <li><b>frames per second</b> : 25.641</li> <li><b>video micrometers per pixel</b> : 4.20853</li> <li><b>number of segmented skeletons</b> : 18752</li> </ul>
LX982 ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V | 2010-03-19T11:02:41+00:00
<blockquote> <p>This experiment is part of the <em>C.elegans behavioural database</em>. For more information and the complete collection of experiments visit http://movement.openworm.org</p> </blockquote> <ul> <li><b>preview link</b> : https://www.youtube.com/watch?v=ZIE5kolDTeQ</li> <li><b>strain</b> : LX982</li> <li><b>timestamp</b> : 2010-03-19T11:02:41+00:00</li> <li><b>gene</b> : ocr-1;ocr-2;ocr-4</li> <li><b>chromosome</b> : IV;V</li> <li><b>allele</b> : ok132;ak47;vs137</li> <li><b>strain_description</b> : ocr-4(vs137)ocr-2(ak47)IV; ocr-1(ok132)V</li> <li><b>sex</b> : hermaphrodite</li> <li><b>stage</b> : adult</li> <li><b>ventral_side</b> : clockwise</li> <li><b>media</b> : NGM agar low peptone</li> <li><b>arena</b> : <ul> <li><b>style</b> : petri</li> <li><b>size</b> : 35</li> <li><b>orientation</b> : away</li> </ul> </li> <li><b>food</b> : OP50</li> <li><b>habituation</b> : 30m wait</li> <li><b>who</b> : Laura Grundy</li> <li><b>protocol</b> : Method in E. Yemini et al. doi:10.1038/nmeth.2560. Worm transferred to arena 30 minutes before recording starts.</li> <li><b>lab</b> : <ul> <li><b>name</b> : William R Schafer</li> <li><b>location</b> : MRC Laboratory of Molecular Biology, Hills Road, Cambridge, CB2 0QH, UK</li> </ul> </li> <li><b>software</b> : <ul> <li><b>name</b> : tierpsy (https://github.com/ver228/tierpsy-tracker)</li> <li><b>version</b> : cbfc23eb4f1ac2f29be75ade7a937eed58a5b219</li> <li><b>featureID</b> : @OMG</li> </ul> </li> <li><b>base_name</b> : ocr-4;2;1 on food L_2010_03_19__11_02_41___8___7</li> <li><b>total time (s)</b> : 899.303</li> <li><b>frames per second</b> : 25.9067</li> <li><b>video micrometers per pixel</b> : 4.20853</li> <li><b>number of segmented skeletons</b> : 18691</li> </ul>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.