Changes for corpus-construction/which-lemmas-does-the-corpus-have/pos-per-lemmas-table.ipynb: 152 added lines, 47 removed lines.
Original line number
Diff line number
Diff line
{
"cells": [
{
"cell_type": "markdown",
"metadata": {},
"source": [
"# Which lemmas does the corpus have?\n",
"\n",
"This script was written for correction purposes: it is meant to display the lemmas present inside the corpus and their associated forms and POS, so as to assess the coherence of the linguistic encoding.\n",
"\n",
"Note: This script may be used on another TEI-XML corpus, if lemma/POS/regularization is structured the same way.\n",
"### FUNCTION: extract the modernized text from one token\n",
"\n",
"Note: As it is widely used throughout my scripts, this function may be made into a separate Python file in the future, to be imported in other scripts."
]
},
{
"cell_type": "code",
"execution_count": 1,
@@ -7,55 +42,61 @@
"outputs": [],
"source": [
"def extraire_forme(word):\n",
" # et on crée la chaîne \"texte\", pour l'instant vide.\n",
" texte = \"\"\n",
" \n",
" # Si <w> n'a pas d'enfant, on récupère le texte tel-quel.\n",