Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
4
datasets available to search
ShareScore release 0.9.0
Dataset results
4 results for “Alsatian Dialects”
Pronunciation Dictionaries for the Alsatian Dialects
<p>This dataset contains a collection of pronunciation dictionaries which were manually transcribed using the X-SAMPA transcription system. The transcriptions were performed based on audio recordings available on the following websites :</p> <ul> <li>OLCA: <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Elsässich Web diktionnair: <a href="http://www.ami-hebdo.com/elsadico/index.php">http://www.ami-hebdo.com/elsadico/index.php</a></li> </ul> <p>The dataset was produced in the context of the RESTAURE project, funded by the French ANR. The transcription process is described in the following research report : 10.5281/zenodo.1174219. It is also detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704814">http://hal.archives-ouvertes.fr/hal-01704814</a></p> <p>Three pronunciation dictionaries are available :</p> <ul> <li>elsassich_dico.csv : 702 transcriptions from the "Elsässich Web diktionnair"</li> <li>olca67.csv : 1,458 transcriptions from the "OLCA" lexicons for the northern part of the Alsace region (Bas-Rhin)</li> <li>olca68.csv : 1,401 transcriptions from the "OLCA" lexicons for the southern part of the Alsace region (Haut-Rhin)</li> </ul>
Lexicon of Place Names in the Alsatian Dialects
<p>This dataset contains a lexicon of place names in the Alsatian dialects. These place names were collected from several resources and manually categorised according to location types defined in the QUAERO project:</p> <ul> <li>loc.fac – Facility</li> <li>loc.phys.astro – Astronym</li> <li>loc.phys.geo – Geonym</li> <li>loc.phys.hydro – Hydronym</li> <li>loc.adm.nat – Country</li> <li>loc.adm.reg – Region</li> <li>loc.adm.sup – Supranational</li> <li>loc.adm.town – City</li> <li>loc.oro – Odonym</li> </ul> <p>The CSV file contains 4 columns:<br> 1. Place name in Alsatian<br> 2. Place name in French<br> 3. Quaero category<br> 4. Source(s): WikiAls (articles from the Alemannic Wikipedia), WikiFr (articles from the French Wikipedia), corpus (Wikipedia articles from the Alemannic Wikipedia and chronicles from an information magazine published by the Haut-Rhin department (southern Alsace) General Council). This field also indicates whether the same spelling can be found in other lexicons of place names for Alsatian: AlsaDico (Edmond Jung. <em>L’alsadico : 22 000 mots et expressions français-alsacien</em>. La<br> Nuée bleue, Strasbourg, 2006.) and Elsàsser (Marc Hug. <em>Toponymes d’Alsace</em>. Online, <a href="http://elsasser.free.fr/NomCommu/ecrantot.html">http://elsasser.free.fr/<br> NomCommu/ecrantot.html</a>, 2007.)</p> <p>The dataset was produced in the context of the RESTAURE project, funded by the French ANR. The lexicon is also decribed in the following article: <a href="https://hal.archives-ouvertes.fr/hal-01702656">https://hal.archives-ouvertes.fr/hal-01702656</a>.</p>
Annotated Corpus for the Alsatian Dialects
<p>This corpus contains a collection of texts in the Alsatian dialects which were manually annotated with parts-of-speech, lemmas, translations into French and location entities.</p><p>The corpus was produced in the context of the RESTAURE project, funded by the French ANR. The current version of the corpus contains 21 documents and 12,907 syntactic words. The annotation process is detailed in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01704806">http://hal.archives-ouvertes.fr/hal-01704806</a></p><p><strong>Information about version 3</strong></p><p>Version 3 corrects some minor errors in the CONLL-U files: wrong token indexes after multiword tokens and missing _ in glosses. In addition, all files are concatenated into a single CONLL-U file.</p><p><strong>Information about version 2</strong></p><p>Version 2 contains the same annotated documents as version 1, but some errors have been corrected and the annotated corpus is provided in the <a href="http://universaldependencies.org/format.html">CoNLL-U format</a></p><p>The untokenised and unannotated versions of the documents are found in the "txt" folder. The annotated versions of the documents are found in the "ud" folder (<a href="http://universaldependencies.org/format.html">CoNLL-U format</a>).</p><p>In addition to the form, the lemma and the part-of-speech additional information is also provided:</p><ul><li>translation of the lemma into French (Gloss field)</li><li>annotation of location names (NamedType field)</li></ul>
Tesseract OCR models for the Alsatian dialects
<p>This dataset provides trained Tesseract (<a href="https://github.com/tesseract-ocr/tesseract">https://github.com/tesseract-ocr/tesseract</a>) OCR models for the Alsatian dialects. These models were developed in the context of the RESTAURE project, funded by the French ANR. </p> <p>Two models are provided :</p> <p>The first model, ISKO_2015, has been presented in the following article: <a href="http://hal.archives-ouvertes.fr/hal-01252241">https://hal.archives-ouvertes.fr/hal-01252241</a>. The Tesseract model has been trained using the jTessBoxEditor tool (<a href="http://vietocr.sourceforge.net/training.html">http://vietocr.sourceforge.net/training.html</a>), Version 1.4 (2 May 2015), based on images automatically generated from the training texts (excerpts from 7 different printed works, totalling about 9,000 words). The generation of the images used a 36pt font size, and two fonts were used (Arial and Times New Roman), with their normal and italic variants.<br> The Tesseract model (gsw.traineddata) can be used with Tesseract 3.0x.</p> <p>The second model, 2018, has been trained for Tesseract 4.0x, using jTessBoxEditor version 2.0.1 (28 July 2018). Again, images were automatically generated from the training text. The training text is different from the one used for the ISKO_2015 model and is "artificial", in the sense that it has been built by appending word n-grams extracted from a large variety of published texts in Alsatian, for a time period spanning 2 centuries and for different text genres. The images corresponding to this training text have been automatically generated with the Tesseract text2image tool, using the following parameters: --ptsize=36 --leading=20. The fonts used are listed in the gsw.font_properties file.</p> <p>Dictionary data has also been used for training. We conflated Alsatian words found in several lexicons and corpora:</p> <ul> <li>Lexicons produced by the OLCA (Office pour la Langue et les Cultures d'Alsace et de Moselle): <a href="http://www.olcalsace.org/fr/lexiques">http://www.olcalsace.org/fr/lexiques</a></li> <li>Lexicon from a Wiktionary user page: <a href="http://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais">https://fr.wiktionary.org/wiki/Utilisateur:Laurent_Bouvier/alsacien-fran%C3%A7ais</a></li> <li>Lexicon from the ACPA association: <a href="http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm">http://web.archive.org/web/20160302234127/http:/culture.alsace.pagesperso-orange.fr/dictionnaire_alsacien.htm</a></li> <li>Chronicles published by Raymond Matzen in the local newspaper "Les Dernières Nouvelles d'Alsace"</li> <li>Transcriptions of television shows found in Erhart, P. (2012). <em>Les dialectes dans les médias: quelle image de l’Alsace véhiculent-ils dans les émissions de la télévision régionale?</em>, Université de Strasbourg, <a href="http://www.theses.fr/167563386">http://www.theses.fr/167563386</a></li> <li>French-Alsatian parallel corpus provided by the OLCA</li> <li>Excerpts from Adolf, P. (2006). <em>Dictionnaire comparatif multilingue: français-allemand-alsacien-anglais.</em>, Strasbourg, France, Midgard, 2006, 373 p.</li> </ul> <p>The Tesseract models can be used for instance using the gImageReader tool (<a href="https://github.com/manisandro/gImageReader">https://github.com/manisandro/gImageReader</a>), which provides a graphical user interface for the Tesseract tool. </p> <p>When evaluated against the same test corpus (prose by Marie Hart, theater and poetry by Gustave Stokopf and prose by Charles Zumstein, totalling about 4,900 words), both models achieve roughly the same performance levels. Usually, even better performance levels can be achieved by combining the Alsatian-specific model with the French and German models available for Tesseract (available from <a href="https://github.com/tesseract-ocr/tessdata">https://github.com/tesseract-ocr/tessdata</a>)</p>
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.