- La version Gold du corpus MICLE <https://www.unicaen.fr/projet_de_recherche/micle/> est composée de 14 textes français (de Normandie et anglo-normand) et de 3 textes en ancien vénitien. Chaque texte a son propre dossier du type "DATE_Nom", dans lequel se trouve différents sous-dossiers par format d'encodage.
- Les textes ont été sélectionnés par les porteurs du projet (Pierre Larrivée et Cécilia Poletto), aidés par Natalia Romanova et Francesco Pinzin. Leah Pavcic and Francesca Santangelo ont également contribué à l'édition numérique.
*Annotation de la partie française*
- La partie française du corpus a été annotée automatiquement par HOPS <https://github.com/hopsparser/hopsparser> (Grobl & Benoît, 2021, https://hal.archives-ouvertes.fr/hal-03223424/file/HOPS_final.pdf). L'annotation a été corrigée manuellement pour les UPOS, et les lemmes ont été annotées grâce aux dictionnaires PRESTO <https://presto.ens-lyon.fr/> et AND <https://anglo-norman.net/>. L'étiquetage UD (<https://universaldependencies.org/format.html>) a été ensuite converti dans les formats PRESTO et UPENN <https://www.ling.upenn.edu/hist-corpora/annotation/index.html>. Les outils ayant servi à la conversion sont disponibles dans le dossier correspondant. Les trois jeux d'étiquettes (UD, UPENN, PRESTO) autorisent différents degrés d'analyse et leur combinaison permet d'affiner les résultats des questions de recherche.
- L'annotation a été faite par Mathieu Goux, avec l'aide de Natalia Romanova (lemmatisation, parties du discours et fonctions syntaxiques.)
*Annotation de la partie vénitienne*
- La partie vénétienne a été annotée manuellement, en l'absence de documentation suffisante pour cette langue, directement dans le système UPENN. Les étiquettes ont ensuite été converties selon le modèle UD pour servir de futur modèle d'entraînement, avec quelques indications syntaxiques.
- L'annotation a été faite par Francesco Pinzin (parties du discours et fonctions syntaxiques), avec une aide substantielle de Leah Pavcic pour l'annotation en parties du discours.
*Édition numérique*
- Pour l'édition numérique des textes, les choix de transcriptions et de découpages des phrases, merci de vous rendre sur la page de documentation générale sur le site du projet : <https://www.unicaen.fr/projet_de_recherche/micle/> ou sur le portail TXM-Crisco <https://txm-crisco.huma-num.fr/txm/>.
@@ -50,14 +56,20 @@
- The Gold version of the MICLE <https://www.unicaen.fr/projet_de_recherche/micle/> corpus is composed of 14 French texts (from Normandy and Anglo-Norman) and 3 Old Venetian texts. Each text has its own folder of type "DATE_Nom", in which there are different subfolders by encoding format.
- The texts were selected by the project leaders (Pierre Larrivée and Cécilia Poletto), assisted by Natalia Romanova and Francesco Pinzin. Leah Pavcic and Francesca Santangelo also contributed to the digital edition.
*French annotation*
- The French part of the corpus was automatically annotated by HOPS <https://github.com/hopsparser/hopsparser> (Grobl & Benoît, 2021, https://hal.archives-ouvertes.fr/hal-03223424/file/HOPS_final.pdf). The annotation was manually corrected for UPOS, and the lemmas were annotated using PRESTO <https://presto.ens-lyon.fr/> and AND <https://anglo-norman.net/> dictionaries. The UD tagset (<https://universaldependencies.org/format.html>) was then converted into PRESTO and UPENN <https://www.ling.upenn.edu/hist-corpora/annotation/index.html> formats. The tools used for the conversion process will be available in the corresponding folder. The three tagsets (UD, UPenn and Presto) are used for the French corpus as each offers a slightly different level of analysis and combining them permits finetuning queries depending on the research question.
- The annotation was done by Mathieu Goux, with the help of Natalia Romanova (lemmatisation, parts of speech and syntactic functions).
*Venetian annotation*
- The Venetian part was annotated manually, in the absence of sufficient documentation for this language, directly in the UPENN system. The labels were then converted to the UD model to serve as a future training model, with some syntactic indications.
- The annotation was done by Francesco Pinzin (parts of speech and syntactic functions), with substantial help from Leah Pavcic for the part-of-speech annotation.
*Digital edition*
- For digital editing of the texts, choices of transcriptions and sentence breakdowns, please check the documentation page on the main website: <https://www.unicaen.fr/projet_de_recherche/micle/>.