Find research datasets worth reusing
Search datasets from major research repositories and use ShareScore to quickly assess how well each record supports discovery, access, and reuse.
31
datasets available to search
ShareScore release 0.9.0
Dataset results
31 results for “French language”
Polifonia Corpus - Books Module Metadata - French Language (Full)
<p>We release the Metadata of the Books module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Linguistic praxeological organization of French as the schooling language
<p>Elements of the linguistic praxeological organization of French as the schooling language used to produce the framework (https://zenodo.org/deposit/4462850) used for the PEAPL project (https://blog.hepfr.ch/create/peapl/) and the framework used for the COMPER project (https://comper.fr/accueil).</p> <p>The following article will give explanations about the elaboration of the linguistic praxeological organizations: https://www.researchgate.net/publication/333244876_Francais_langue_de_scolarisation_Reflexions_sur_les_referentiels_de_competences_et_l'adaptive_learning</p>
French-language art critics bibliographies dataset
<p>French-language art critics active between the mid-19<sup>th</sup> and 20<sup>th</sup> centuries (see <a href="http://critiquesdart.univ-paris1.fr/en">http://critiquesdart.univ-paris1.fr/en</a>). Each author’s page displays their primary and secondary bibliographies and identified archival sources. These documents are based on extensive research to produce primary bibliographies that can be considered to be comprehensive. For those writers whose complete works have been published, only their writing on art is listed. This online directory doubles as a database of primary bibliographies for the critics listed.</p> <p>This dataset is neutral and non-discriminating, with no typology or classification other than the text format. It is multidisciplinary and makes no claim to assess the critical value of a text, preferring to provide as comprehensive an overview as possible of the authors’ literary output in whatever field (literature, art, politics, history, etc.) or medium (photography, film, fine arts, architecture, etc.) their interest encompassed.</p> <p>Since these authors all wrote in French, they come from the artistic scenes in France, Belgium and Switzerland, not to mention the colonies of the period. Authors who wrote in more than one language are also included.</p> <p>Our intention is to showcase research into art criticism, improve access to the documents and create links between researchers. The site will be gradually updated and enriched by the addition of further authors and greater detail for its current bibliographies.</p> <p><strong>Full mysql database : "<a href="https://zenodo.org/api/files/62e9c484-74ae-456f-91f1-bb8dde962d11/critiquesdart_08_02_2022.sql?versionId=02fdf2bb-4b38-43b0-9c2c-1bf32b8997a9">critiquesdart_08_02_2022.sql</a>", see </strong><a href="https://doi.org/10.7202/1079443ar">https://doi.org/10.7202/1079443ar</a> for more details (in french).</p> <p><strong>Data set 1 : " <a href="https://zenodo.org/record/3999374/files/database_journal_list.csv?download=1">database_journal_list.csv</a> " description (en/fr) :</strong></p> <p><em><strong>description of the fields:</strong></em></p> <ul> <li>"title": title of the journal</li> <li>"ISSN": International identifier for serial publications</li> <li>"temporal_cover": Start and end dates of publication</li> <li>"city": place of publication, if known, otherwise NULL</li> </ul> <p><strong><em>descriptif des champs :</em></strong></p> <ul> <li>"titre": titre de la revue</li> <li>"ISSN": Identifiant International des publications en série</li> <li>"couverture_temporelle": Dates de début et de fin de publication</li> <li>"ville": lieu de publication, si connu, sinon NULL</li> </ul> <p><strong>Dataset 2 : " <a href="https://zenodo.org/api/files/4023bb1d-b185-432c-bbae-1d4b27277468/Notices_critiques.csv?versionId=d682a60a-3e7e-4dc3-aee0-14636b4ac387">Notices_critiques.csv </a> " description (en/fr) : list of all the critics in the database (see http://critiquesdart.univ-paris1.fr/annuaire_critiques.php).</strong></p> <p><em><strong>description of the fields:</strong></em></p> <ul> <li>"id": identifier in the database</li> <li>"first name": First name of the author, if known, otherwise empty</li> <li>"name": Author's last name</li> <li>"birth": Year of birth</li> <li>"death": date of death</li> <li>"ISNI": International Standard Name Identifying people</li> </ul> <p><em><strong>descriptif des champs :</strong></em></p> <ul> <li>"id ": identifiant dans la base</li> <li>"prenom ": Prénom de l'auteur, si connu, sinon vide</li> <li>"nom": Nom de famille de l'auteur</li> <li>"naissance": Année de naissance</li> <li>"mort": date de mort</li> <li>"ISNI": International Standard Name Identifier des personnes</li> </ul> <p><strong>Dataset 3 : "</strong><a href="https://zenodo.org/api/files/4023bb1d-b185-432c-bbae-1d4b27277468/full_notices_dataset.csv?versionId=8d851f6a-f0ba-4989-ba85-a5b825fbd164">full_notices_dataset.csv</a>" <strong>description (en/fr) :</strong> List of all the notices</p> <p><em><strong>description of the fields:</strong></em></p> <ul> <li>"critical_id": id of the review in the database</li> <li>"title": title of the review</li> <li>"subtitle": subtitle of the review, otherwise NULL</li> <li>"type": type of criticism</li> <li>"author_id": author's id in the database</li> <li>"name": name of the author reviewer</li> <li>"first name": first name of the author critic</li> <li>"annee_ouvrage": year of publication if the review is a work or part of a work, otherwise NULL</li> <li>"coordinator": potential coordinator of the structure, otherwise NULL</li> <li>"titre_ouvrage": title of the work if the review is an article, otherwise NULL</li> <li>"annee_periodique *": year of publication if the review is an article, otherwise NULL</li> <li>"periodic_title": name of the journal if the review is an article, otherwise NULL</li> <li>"city": place of publication, if known, otherwise NULL</li> </ul> <p><em><strong>descriptif des champs :</strong></em></p> <ul> <li>"id_critique": id de la critique dans la base</li> <li>"titre": titre de la critique</li> <li>"sous titre": sous titre de la critique, sinon NULL</li> <li>"type": type de critique</li> <li>"id_auteur": id de l'auteur dans la base</li> <li>"nom": nom du critique auteur</li> <li>"prenom": prénom du critique auteur</li> <li>"annee_ouvrage": année de parution si la critique est un ouvrage ou partie d'un ouvrage, sinon NULL</li> <li>"coordonateur": coordonateur potentiel de l'ouvrage, sinon NULL</li> <li>"titre_ouvrage": titre de l'ouvrage si la critique est un article, sinon NULL</li> <li>"annee_periodique *": année de parution si la critique est un article, sinon NULL</li> <li>"titre_periodique": nom de la revue si la critique est un article, sinon NULL</li> <li>"ville": lieu de publication, si connu, sinon NULL</li> </ul>
Polifonia Corpus - Periodicals Module Metadata - French Language
<p>We release the Metadata of the Periodicals module of the Polifonia Textual Corpus. According to the availability from the source origin, the Metadata may include the URL from which a text of the Books corpus is accessible, along with the title, the author, the year of publication, and the publisher. Metadata allows for a complete reconstruction of the corpus as we cannot make the actual texts available because they are subject to heterogeneous licensing.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Extract from the PEAPL framework : Modelling "Writing sentences" competency (French as the schooling language)
<p>Competences, skills and knowledges that make up "writing sentences" competency, based on the linguistic praxeological organization of French as the schooling language (https://doi.org/10.5281/zenodo.4001381). This is an extract of the general framework, some of the visible objects are linked to other objects in other main competences. Orange links show how pedagogic ressources (game levels) are linked to framework objects.</p> <p>This framework is used for the PEAPL (peapl.eu) project to link activities in the GamesHub platform (for example activities using "L'Orthodyssée des Gram" grammar online game, available at https://www.lafamillegram.ch/#)</p> <p> </p> <p> </p> <p> </p>
Polifonia Corpus - Encyclopedic Module Metadata - French Language
<p>We make available the Metadata related to the Wikipedia pages that constitute the Encyclopedic Module of the Polifonia Textual Corpus. Metadata for this module includes, per each Wikipedia page, its Wikipedia ID, BabelNet ID, gloss, resource type (that can be named entity or concept), Lemmata, Sensekey, WikiData ID.</p> <p>Full description at <a href="http://github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Annotated corpus samples of resultative nomination constructions in 4 languages (Dutch, English, French, Spanish)
<p>This dataset includes an annotated corpus sample retrieved from Sketch Engine of 3200 (i.e. 800x4) resultative constructions in 4 languages (Dutch, English, French, Spanish) with 16 (i.e. 4x4) different nomination verbs (Dutch: kronen 'crown', verkiezen 'elect', promoveren 'promote', uitroepen 'proclaim; English: crown, elect, promote, proclaim; French: couronner 'crown', élire 'elect', promouvoir 'promote', proclamer 'proclaim'; Spanish: coronar 'crown', elegir 'elect', ascender 'promote', proclamar 'proclaim').</p>
PxCorpus : A Spoken Drug Prescription Dataset in French for Spoken Language Understanding and Dialogue
<h2><strong>PxCorpus : A Spoken Drug Prescription Dataset in French</strong></h2><p> </p><p>PxCorpus is to the best of our knowledge, the first spoken medical drug prescriptions corpus to be distributed.</p><p>It contains 4 hours of transcribed and annotated dialogues of drug prescriptions in French acquired through an experiment with 55 participants experts and non-experts in drug prescriptions.</p><p> </p><p>The automatic transcriptions were verified by human effort and aligned with semantic labels to allow training of NLP models. The data acquisition protocol was reviewed by medical experts and permit free distribution without breach of privacy and regulation.</p><p> </p><h3>Overview of the Corpus</h3><p>The experiment has been performed in wild conditions with naive participants and medical experts.</p><p>In total, the dataset includes 2067 recordings of 55 participants (38% non-experts, 25% doctors, 36% medical practitioners), manually transcribed and semantically annotated.</p><p> </p><p>| Category | Sessions | Recordings | Time(m)|</p><p>|-----------------------| ------------- | --------------- | ----------- |</p><p>| Medical experts | 258 | 434 | 94.83 |</p><p>| Doctors | 230 | 570 | 105.21 |</p><p>| Non experts | 415 | 977 | 62.13 |</p><p>| Total | 903 | 1981 | 262.27 |</p><p> </p><h3>License</h3><p>We hope that that the community will be able to benefit from the dataset which is distributed with an attribution 4.0 International (CC BY 4.0) Creative Commons licence.</p><p> </p><h3>How to cite this corpus</h3><p> </p><p>If you use the corpus or need more details please refer to the following paper: A spoken drug prescription datset in French for spoken Language Understanding</p><p> </p><p>@InProceedings{Kocabiyikoglu2022,</p><p> author = "Alican Kocabiyikoglu and Fran{\c c}ois Portet and Prudence Gibert and Hervé Blanchon and Jean-Marc Babouchkine and Gaëtan Gavazzi",</p><p> title = "A spoken drug prescription datset in French for spoken Language Understanding",</p><p> booktitle = "13th Language Ressources and Evaluation Conference (LREC 2022)",</p><p> year = "2022",</p><p> location = "Marseille, France"</p><p>}</p><p>a more complete description of the corpus acquisition is available on arxiv </p><p>@misc{kocabiyikoglu2023spoken,</p><p> title={Spoken Dialogue System for Medical Prescription Acquisition on Smartphone: Development, Corpus and Evaluation},</p><p> author={Ali Can Kocabiyikoglu and François Portet and Jean-Marc Babouchkine and Prudence Gibert and Hervé Blanchon and Gaëtan Gavazzi},</p><p> year={2023},</p><p> eprint={2311.03510},</p><p> archivePrefix={arXiv},</p><p> primaryClass={cs.CL}</p><p>}</p><p> </p><h3>Project Structure</h3><p> </p><p>The project contains the following elements</p><p>.</p><p>├── LICENSE</p><p>├── PxDialogue/</p><p>├── PxSLU/</p><p>├── readme.md</p><p> </p><h3>PxSLU : Prescription Corpus for Spoken Language Understanding</h3><p> </p><h4>Directory Structure</h4><p>.</p><p>├── LICENSE</p><p>├── metadata.txt</p><p>├── paths.txt</p><p>├── PxSLU_conll.txt</p><p>├── readme.md</p><p>├── recordings</p><p>├── seq.in</p><p>├── seq.label</p><p>├── seq.out</p><p>├── Demo.ipynb</p><p>└── verifications.py</p><p> </p><h4>Recordings</h4><p> </p><p>The recordings directory contains the 903 recording sessions. Each session can contain several recordings. For instance,</p><p>the directory</p><p> recordings/J7aVvWb67L</p><p>contains the records </p><p> recording_0.wav recording_2.wav</p><p>which represent two attempts to record a drug prescription</p><p> </p><p>All records are stored as mono channel wav files of 16kHz 16bits signed PCM</p><p> </p><h4>Paths</h4><p> </p><p>contains the list of all the .wav files in the recordings directory</p><p>00MYcyVK0t/recording_0.wav</p><p>00MYcyVK0t/recording_2.wav</p><p>02Qp6ICj9Q/recording_0.wav</p><p>02Qp6ICj9Q/recording_1.wav</p><p>...</p><p> </p><p>All other files (metadata.txt, seq.*) refer to this list to describe the recording.</p><p> </p><h4>Metadata</h4><p> </p><p>contains the information about the participants:</p><p> </p><p>48,60+,F,non-expert</p><p>48,60+,F,non-expert</p><p>24,18–28,F,doctor</p><p>24,18–28,F,doctor</p><p>...</p><p> </p><p>The first column is the participant unique id, the second is the age range, the third is the gender and the final is the category of the participant in {doctor,expert, non-expert}. doctor correspond to a physician, (other)expert to a pharmacist or a biologist specialized in drugs while non-expert are other people not entering in these categories. The lines are synchronised with the paths.txt lines.</p><p> </p><h4>Labels</h4><p> </p><p>the three files seq.label, seq.in, seq.out represent respectivly the intent, the transcript and the entities in BIO format.</p><p> </p><p> seq.label | seq.in | seq.out</p><p>medical_prescription | flagyl 500 milligrammes euh qu/ en... | B-drug B-d_dos_val B-d_dos_up O O ...</p><p>medical_prescription | 3 comprimés par jour matin midi ... | B-dos_val B-dos_uf O O B-rhythm_tdte B-rhythm_tdte O B-rhythm_tdte ...</p><p> ... | ... | ...</p><p> </p><p>These lines are synchronised with the paths.txt lines.</p><p> </p><p>Another file "PxSLU_conll.txt" is provided in a format inspired by the conll format (https://universaldependencies.org/format.html). However, this one is *not* aligned with the acoustic records file paths.txt.</p><p> </p><h4>Scripts</h4><p> </p><p>verifications.py performs the checking of the alignement of all the seq.* paths.txt and metadata files. A user of the dataset does not need to use this script unless she plan to extend the datasets with her own data.</p><p>Demo.ipynb is a jupyter notebook that a user can run to search through the dataset. It is intended to let the user have a quicker and smoother view on the dataset.</p><p> </p><p> </p><h4>Data splits</h4><p>In the data_splits folder, you can find a data split of this dataset organized as following:</p><p>- train.txt: medical experts + non experts (80%) = 1128 samples</p><p>- dev.txt: medical experts + non experts (20%) = 283 samples</p><p>- test.txt: doctors (100%) = 570 samples</p><p> </p><p>Each file contains references to line numbers of the corpus. For example, first line of the test.txt is 904, seq.in file contains the utterance "nicopatch". Users can access the labels, slots, metadata using the same line number 904 in the parallel files (paths.txt,seq.out,seq.label,...).</p><p> </p><h3>PxDialogue : Prescription recording corpus for dialogue systems</h3><p> </p><p>PxDialogue corpus comes as an extension of the PxSLU corpus and provides additional information about the dialogues that was collected through spoken dialogue. This corpus includes two additional files:</p><p>├── events.txt</p><p>├── dialogue_annotations.txt</p><p> </p><p> </p><h4>Events.txt:</h4><p> </p><p>For each dialogue session, all dialogue events are given in this text file</p><p>which can be used to train/evaluate dialogue systems.</p><p> </p><p>Usage example:</p><p> </p><p>PxSLU (paths.txt)</p><p>- 00MYcyVK0t/recording_0.wav</p><p>- 00MYcyVK0t/recording_2.wav</p><p> </p><p>PxDialogue (events.txt)</p><p>- (-1, 'START', 'APP', None, 0) (1, 'user', 'ASR', 'flagyl 500 mg en cachet pendant 8 jours', 30) (1, 'system', 'TTS', 'Choisissez le médicament correspondant à votre recherche', 34) (2, 'user', 'UI', 'listview_item_clicked', 40) (2, 'system', 'TTS', 'Pourriez vous préciser la posologie pour le patient?', 40) (3, 'user', 'ASR', '3 comprimés par jour matin midi et soir pendant 10 jours', 64) (3, 'system', 'TTS', "Est-ce que vous confirmez l'ajout de cette prescription sur la liste?", 66) (4, 'user', 'UI', '/inform{"validate":"validate"}', 73) (4, 'system', 'TTS', 'Prescription validée avec succès. Traitement ajouté sur le dossier du patient', 73) (-1, 'END', 'APP', '', 73)</p><p>- N/A</p><p> </p><p>For example, in this dialogue session (00MYcyVK0t), there are two recordings.</p><p>The events are given in a single row for each dialogue session once in the</p><p>first recording (recording_0). Dialogues are described in form of events</p><p>where each action taken by the user or the system is considered as a dialogue</p><p>turn in a tuple form.</p><p> </p><p>(-1, 'START', 'APP', None, 0)</p><p> </p><p>- First element of the event is the dialogue turn number. -1 means that the application</p><p>is initialized.</p><p>- Second element describes who initiated the event: user, system, START, END</p><p>- Third element describes the type of the event: APP (start and end events)</p><p>, ASR (automatic speech recognition), TTS (text-to-speech), UI (user interface)</p><p>User clicks on buttons triggers sometimes explicit intent recognition. For</p><p>example (4, 'user', 'UI', '/inform{"validate":"validate"}') describes the</p><p>explicit intent of validation of the prescription.</p><p>- Fourth element is the timestamp (in seconds)</p><p> </p><h4>Dialogue annotations</h4><p> </p><p>We also include a manual annotation for dialogues (dialogue_annotations.txt) which indicates for each recording, if the system gave the correct answer given the utterance.</p><p> </p><p>Each line contains a keyword, either [Fail] or [OK]. The following example shows a dialogue sample with annotations:</p><p> </p><p>| dialogue_annotations.txt | paths.txt | seq.in |</p><p>|----------------------------------|----------------------------------------|--------------------------------------------------------------------------------------|</p><p>| OK | 14yHtAe555/recording_0.wav | oxytetracycline solution euh |</p><p>| OK | 14yHtAe555/recording_1.wav | oxytetracycline solution 5 gouttes matin et soir pendant 14 jours |</p><p>| Fail | 14yHtAe555/recording_2.wav | oxytetracycline solution 5 gouttes matin et soir pendant 14 jours |</p><p> </p><p>For these 3 dialogues, the dialogue annotations are accordingly OK, OK and Fail.</p><p>[Ok] means that the dialogue system reacted correctly to the input.</p><p>[Fail] means that the action of the system after this utterance should not be used</p><p>for evaluation or training. </p><p>We can notice that in the first utterance, the information are missing, however</p><p>after the second example the system normally have all of the required slots</p><p>for the prescription validation.</p><p> </p><p>It is to note that free comments added using the ASR system were noted as</p><p>Fail as these dialogues did not enter the dialogue state tracking. In this</p><p>example, the last utterance is recorded as a free comment by the prescriber</p><p>and was annotated as Fail.</p><p> </p><h4>Linking audio records to ASR events</h4><p>In order to link audio records to ASR events, the user has to use both paths.txt and events.txt</p><p> </p><p>For example for the following dialogue session (lines 1:2 of paths.txt):</p><p>00MYcyVK0t/recording_0.wav</p><p>00MYcyVK0t/recording_2.wav</p><p> </p><p>Events.txt include two ASR events:</p><p>(1, 'user', 'ASR', 'flagyl 500 mg en cachet pendant 8 jours', 30)</p><p>(3, 'user', 'ASR', '3 comprimés par jour matin midi et soir pendant 10 jours', 64)</p><p> </p><p>These ASR events corresponds to the recording files that can be found in the recordings folder.</p><p> </p><h4>Available user action annotations</h4><p> </p><p>Events.txt include annotations such as below with the following explanation:</p><p> </p><p>- /inform{"validate":"validate"} : User clicks on the validate button after seeing the prescription</p><p>- /inform{"validate":"refuse"} : User clicks on the refuse button after seeing the prescription</p><p>- ASR : User clicks on the push-to-talk button to record an utterance</p><p>- listview_item_clicked : User clicks on the list to choose a drug</p><p>- listview_cancel_clicked : User clicks on the cancel button after seeing a list of drugs</p><p>- FREE_COMMENT_ADDED : User clicks and records a free-form utterance by clicking "add free comment" button</p><p>- EMPTY_UTTERANCE: Recording containing an empty utterance</p><p>- APP_CRASH : An application crash that happened in the dialogue turn</p><p>- EVAL_FINISH_APPROVED : User clicks on the final upload button to finish the experiment</p><p>- RESTART_CONVERSATION_SESSION : User clicks on the restart conversation button</p><p>- RESTART_CANCEL_CLICKED : User cancels the restart process by clicking on the cancel button</p><p>- EVAL_FINISH_CANCELED : User cancels the final upload process by clicking on the cancel button</p><p> </p><p>** Free Comments: **</p><p>The users had the possibility of recording a speech-to-text message upon viewing a prescription.</p><p>These messages had not beed added to the dialogue state tracking but were visualized on the interface and saved in</p><p>the database. Users can find free comments by searching for FREE_COMMENT_ADDED events in the events.txt to find out</p><p>about these events.</p>
Polifonia Corpus - Encyclopedic Module Data - French Language
<p>We release the data of the Encyclopedic Module of the Polifonia Textual Corpus (Wikipedia pages), collected by selecting from <a href="http://lcl.uniroma1.it/babeldomains/">BabelNet domains</a> all the <a href="https://www.wikipedia.org">Wikipedia</a> musical pages.</p> <p>Full description at <a href="/github.com/polifonia-project/Polifonia-Corpus">https://github.com/polifonia-project/Polifonia-Corpus</a></p>
Data belonging to the article "Syllabification in Language Contact between French and Vietnamese"
<p>This dataset belongs to an article which will appear in the ICAAL 8 conference proceedings and which is dedicated to the syllabification in two language contact settings between Vietnamese and French. The paper discusses gemination and related phenomena. First, it contributes to the theoretical discussion on borrowing processes in loanwords and presents a model that includes all extant views and also accounts for other language contact situations. In our data analysis, we compare loan data to experimental data, and find similar processes in both scenarios. Syllable boundaries are shifted, consonants are geminated and consonant slots are doubled in certain syllabic environments. The emergence of gemination is especially noteworthy as it is only a marginal phenomenon in French and has not been found to exist in Vietnamese. Results provide evidence for our hypothesis that optimal Vietnamese syllables are closed, as these syllables have the highest frequency in the native lexicon. This goes beyond, but also corresponds to the idea that Vietnamese has a dimoraic minimum, and this paper provides more evidence for it.</p>
Figure S4. Presence of Japanese Kanji ideograms on rubber bales found in 2021 along the Brazilian coastline. This historical evidence from 2021 supports the hypothesis of a new wreck as the 2018 and 2019 event had different language entries from the former French Indo-China (Teixeira et al., 2021).
Open the record for dataset details and reuse information.
Data belonging to the thesis intitled "Prosody in Language Contact. French and Vietnamese"
<p>This dataset belongs to our doctoral thesis intitled "Prosody in Language Contact. French and Vietnamese". The dataset contains recordings of oral data as well as tables of written data, transcriptions and annotations. It also contains metadata about oral and written data. Detailed information about this data and how to understand it can be taken from our thesis.</p> <p>In the thesis, we are dealing with prosodic language contact between the languages Vietnamese and French. Our research is devoted to different language contact situations as well as to both directions of language contact. The starting point of our experimental research is the observation of French loanwords in Vietnamese. In this context, we examine prosodic adaptation patterns that speakers have undertaken to adapt the loanwords to Vietnamese. The focus is on the repair of structures that are illicit in Vietnamese: Consonants in certain positions in the syllable are replaced or deleted, consonant clusters are dissolved by epenthesis or by deletion of one of the two consonants, syllable boundaries are shifted, consonants or consonant slots in certain syllabic structures are doubled, and tones are assigned to syllables according to certain patterns.</p> <p>The aim of our experimental investigations then is to find out whether similar or different patterns occur in an instantaneous situation of language contact. Monolingual speakers of Vietnamese are exposed to French stimuli and asked to reproduce them in three different conditions. We have additionally conducted the same experiment with learners of French whose first language is Vietnamese. The experimental data show many similar patterns to the loanword data. However, the data from monolingual speakers in particular display much more variability. Finally, we reverse the direction of language contact: Native speakers of French are asked to reproduce Vietnamese stimuli. In this case, the question is whether certain patterns can be reversed in the opposite direction, which can partly be observed.</p> <p>With the help of the present work, we can gain a deep understanding of systematicity and variability in borrowing and second language acquisition. We conclude that from a phonological perspective, there are many similarities between the two fields of prosodic language contact, and we suggest, for the phenomena under consideration, to understand some aspects in their complexity in a gradual rather than a categorical way.</p>
Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains: Data
<p>This repository contains the data for the following paper:</p> <p><span>Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Daniel Audibert, Xingyu Liu, Cécile Macaire, Adrien Pupier, Yongxin Zhou, Mathilde Aguiar, Felix E. Herron, Magali Norré, Massih R Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, et al.. 2024. <a href="https://aclanthology.org/2024.lrec-main.827">Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized Domains</a>. In <em>Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)</em>, pages 9463–9476, Torino, Italia. ELRA and ICCL.</span></p> <p>1) pretraining data for the Jargon specialized language models</p> <p>2) ECTHR_FR dataset for text classification in the French legal domain</p> <p> </p>
Annotated corpus sample of resultative constructions in cooking recipes in 4 languages (Dutch, English, French, Spanish)
<p>This dataset contains an annotated corpus sample of 4000 (i.e. 1000 per language) occurrences of resultative constructions (e.g. cut the onion thin, whisk the egg whites to a foam, roll the dough into a ball) in 4 languages (Dutch, English, French, Spanish) retrieved from a comparable multilingual corpus of cooking recipes.</p>
Outcome of Autosomal Dominant Polycystic Kidney Disease Patients on Peritoneal Dialysis: a National Retrospective Study Based on Two French Registries (the French Language Peritoneal Dialysis Registry
ClinicalTrials.gov study NCT03948113. IPD Sharing: UNDECIDED. Countries: 1. Publications: 1.
Validation of the Foot Health Status Questionnaire in French Language
ClinicalTrials.gov study NCT06933953. IPD Sharing: Not stated. Countries: 1. Publications: 4.
Validation in French Language of the Questionnaire EARS
ClinicalTrials.gov study NCT03963440. IPD Sharing: Not stated. Countries: 1. Publications: 5.
ANTHROPONYMS WITH EVALUATIVE MEANING IN THE FRENCH LANGUAGE
Open the record for dataset details and reuse information.
SOME CHARACTERISTICS OF WORD SEMANTIC SPECIFICATIONS IN FRENCH AND UZBEKI LANGUAGES
Open the record for dataset details and reuse information.
THE LEXICAL CHARACTERISTICS OF CANADIAN FRENCH INFLUENCED BY LANGUAGE INTERFERENCE
Open the record for dataset details and reuse information.
ScienceDex guides
Understand access before you commit
These curated guides explain access requirements, typical timelines, costs, and reuse considerations for widely used research datasets.
Allen Brain Atlas
Allen Brain Atlas is an Allen Institute collection of brain map atlases, datasets, APIs, and analysis tools covering mouse, human, and non-human primate brain resources.
Annotated Behaviour and Observability Dataset (ABODe)
ABODe is a University of Edinburgh DataShare dataset for behavior classification in group-housed mice using home-cage video, identities, bounding boxes, ground-plate positions, and annotator labels.
DANDI Archive for NWB datasets
DANDI is a BRAIN Initiative archive for publishing and sharing neurophysiology data, including electrophysiology, optophysiology, and behavioral data packaged as NWB and related standards.
International Brain Laboratory public data
The International Brain Laboratory public data releases expose standardized mouse decision-making experiments, including Neuropixels recordings, widefield calcium imaging, behavior, and session metadata accessed through the ONE API.
OpenNeuro
OpenNeuro is a free, open platform for sharing neuroimaging datasets, with public search, dataset pages, and download paths for web, S3, DataLad, and the OpenNeuro CLI.