Simplemma: fast multilingual lemmatization for Python
Fast, dependency-free lemmatization for 54 languages. Pure Python, no models to download, works offline.
- A 19 MB install with no per-language downloads: 9 of the 54 languages have
- ~1.9M tokens/s (German) and ~3.4M (English) with tokenization
- Tunable RAM footprint: ~175 MB, ~50 MB with
low_memory=True, or ~30 MB per
- 0.91 to 0.97 accuracy for 34 languages, German at 0.97 and English at
- Useful utilities included: script-aware tokenizer, rule-based sentence
Purpose
Lemmatization groups the inflected forms of a word so they can be analysed as a single item, identified by the word's lemma, or dictionary form. Unlike stemming, the output is always a valid linguistic form.
Simplemma provides a simple and multilingual approach to looking for base forms. It needs no morphosyntactic information and processes a raw series of tokens, or a text through its built-in tokenizer. It is not as powerful as full-fledged solutions, but it is generic, easy to install and fast, and its small footprint suits contexts where speed and simplicity matter: low-resource settings, teaching, or a baseline for lemmatization and morphological analysis.
Currently, 54 languages are partly or fully supported (see the list below).
Installation
The current library is written in pure Python with no dependencies:
pip install simplemma
pip install -U simplemmafor updatespip install git+https://github.com/adbar/simplemmafor the cutting-edge versionpip install simplemma[marisa-trie]for the lowest memory usage. For a
low_memory=True (see
Memory usage)
Python 3.10 or later is required: the last version supporting 3.8 and 3.9
is simplemma==1.1.2, and simplemma==1.0.0 for 3.6 and 3.7.
Usage
Quick start
Pick a language and apply it to a single word, to a list of tokens, or to a whole text through the built-in tokenizer:
`` python
>> import simplemma
>> simplemma.lemmatize('masks', lang='en')'mask'
>> mytokens = ['Hier', 'sind', 'Vaccines']
>> [simplemma.lemmatize(t, lang='de') for t in mytokens]['hier', 'sein', 'Vaccines']
>> simplemma.is_known('spaghetti', lang='it')True
>> simplemma.simple_tokenizer('Hier sind Vaccines.')['Hier', 'sind', 'Vaccines', '.']
>> simplemma.text_lemmatizer('Hier sind Vaccines.', lang=('de', 'en'))['hier', 'sein', 'vaccine', '.']
### Chaining languages
Chaining several languages can improve coverage, they are used in
sequence:
python
>> from simplemma import lemmatize
>> lemmatize('Vaccines', lang=('de', 'en'))'vaccine'
>> lemmatize('spaghettis', lang='it')'spaghettis'
>> lemmatize('spaghettis', lang=('it', 'fr'))'spaghetti'
python### Greedier decompositiongreedyFor some languages a greedier decomposition is active by default because it helps to strip affixes. It can be triggered manually by setting the
parameter toTrue, which adds an iteration of the search algorithm and may come closer to stemming than to lemmatization.
>> simplemma.lemmatize('ausgezeichneten', lang='de', greedy=False)'ausgezeichnet' # 1 step: reduction to past participle
>> simplemma.lemmatize('ausgezeichneten', lang='de', greedy=True)'auszeichnen' # 2 steps: further reduction to infinitive verb
<!-- include:usage:end -->
Tokenization
python
>> from simplemma import simple_tokenizer
>> simple_tokenizer('Lorem ipsum dolor sit amet, consectetur elit.')['Lorem', 'ipsum', 'dolor', 'sit', 'amet', ',', 'consectetur', 'elit', '.']
pythonThe tokenizer is script-aware: in-word joiners and combining marks stay inside their word, punctuation becomes its own token, and numbers keep a sensible shape.text_lemmatizer()andlemma_iterator()chain tokenization with lemmatization, lowering sentence-initial capitals before lookup.Sentence splitting
>> from simplemma import split_sentences
>> split_sentences('Das Tor fiel in der 95. Minute. Das Spiel war aus.', lang='de')['Das Tor fiel in der 95. Minute.', 'Das Spiel war aus.']
pythonRule-based, with per-language abbreviation data forcs,de,en,fr,nl,pl,pt. Sentence-boundary F1 is 0.98 to 0.998 on held-out UD corpora.langdetect()Language detection
scores a text against languages of interest;in_target_language()returns the single ratio of recognized tokens.
>> from simplemma import langdetect
>> langdetect('"Exoplaneta, též extrasolární planeta, je planeta obíhající kolem jiné hvězdy."', lang=("cs", "sk"))[('cs', 1.0), ('sk', 0.14285714285714285), ('unk', 0.0)] `
Caveats
The dictionary lookup cannot disambiguate when several lemmas are valid, and
diminutives or rare forms may be absent. Working without morphosyntactic
information sets a hard ceiling — a few points behind trained neural
pipelines, hundreds of times faster. See
Usage in the documentation for
details.
Advanced usage via classes and lower memory usage
Instantiating the classes instead of calling the functions gives more
control: a custom
LemmatizationStrategy, a DictionaryFactory with its own
cache size, or one of the low-memory backends selected by low_memory=True.
See Classes and strategies
for the classes and Memory usage
for the low_memory flag and a comparison of the three dictionary backends.
Supported languages
The following languages are available, identified by their BCP 47
language tag, usually
the ISO 639-1 code.
Available languages (2026-08-04):
The Forms column counts the inflected word forms stored in the
dictionary, while Lemmata counts the distinct base forms they map to
(both in thousands). A large gap between the two reflects rich
morphology rather than a data error.
| Code | Language | Forms (10³) | Lemm. (10³) | Acc. | Comments |
| ---- | -------- | ----------- | ------------| ---- | -------- |
|
ar | Arabic | 298 | 49 | 0.91 | on UD AR-PADT; real-world (unsegmented) input scores ≈0.85 |
| ast | Asturian | 154 | 36 | | |
| bg | Bulgarian | 226 | 26 | 0.89 | on UD BG-BTB |
| ca | Catalan | 641 | 64 | 0.95 | on UD CA-AnCora |
| cs | Czech | 363 | 47 | 0.95 | on UD CS-FicTree |
| cy | Welsh | 402 | 21 | 0.94 | on UD CY-CCG |
| da | Danish | 788 | 117 | 0.95 | on UD DA-DDT, alternative: lemmy |
| de | German | 1,116 | 334 | 0.97 | on UD DE-GSD, see also German-NLP list |
| el | Greek | 250 | 28 | 0.93 | on UD EL-GDT |
| en | English | 182 | 78 | 0.96 | on UD EN-LinES, alternative: LemmInflect |
| enm | Middle English | 43 | 6 | | |
| eo | Esperanto | 191 | 18 | 0.95 | on UD EO-PraGo (test-only treebank, no train split) |
| es | Spanish | 824 | 88 | 0.96 | on UD ES-AnCora |
| et | Estonian | 2,690 | 95 | 0.92 | on UD ET-EWT, low coverage |
| fa | Persian | 47 | 14 | 0.95 | on UD FA-Seraji |
| fi | Finnish | 3,547 | 125 | 0.91 | on UD FI-TDT, see this benchmark |
| fr | French | 250 | 37 | 0.96 | on UD FR-Sequoia |
| ga | Irish | 444 | 48 | 0.92 | on UD GA-IDT |
| gd | Gaelic | 73 | 16 | 0.89 | on UD GD-ARCOSG |
| gl | Galician | 426 | 43 | 0.92 | on UD GL-CTG |
| grc | Ancient Greek | 852 | 22 | 0.86 | on UD GRC-PROIEL (best available; no general-register grc treebank exists), alternative: odyCy |
| gv | Manx | 77 | 14 | 0.92 | on UD GV-Cadhan |
| hbs | Serbo-Croatian | 610 | 49 | 0.90 | on UD HR-SET + SR-SET (token-weighted); Croatian and Serbian lists to be added later |
| he | Hebrew | 105 | 10 | 0.93 | on UD HE-HTB; real-world (unsegmented) input scores ≈0.82 |
| hi | Hindi | 86 | 19 | 0.95 | on UD HI-HDTB |
| hu | Hungarian | 1,763 | 45 | 0.88 | on UD HU-Szeged |
| hy | Armenian | 467 | 17 | 0.92 | on UD HY-BSUT |
| id | Indonesian | 22 | 4 | 0.93 | on UD ID-CSUI |
| is | Icelandic | 210 | 18 | 0.82 | on UD IS-GC |
| it | Italian | 358 | 28 | 0.95 | on UD IT-ISDT |
| ka | Georgian | 448 | 16 | 0.85 | on UD KA-GLC |
| la | Latin | 1,289 | 70 | 0.89 | on UD LA-PROIEL, alternative: LatinCy |
| lb | Luxembourgish | 306 | 79 | | only a <1k-token UD treebank available |
| lt | Lithuanian | 365 | 28 | 0.87 | on UD LT-ALKSNIS |
| lv | Latvian | 178 | 15 | 0.84 | on UD LV-LVTB |
| mk | Macedonian | 546 | 41 | 0.92 | on UD MK-MTB (test-only treebank, no train split) |
| ml | Malayalam | 746 | 64 | 0.69 | on UD ML-UFAL (small test-only treebank, no train split), experimental |
| ms | Malay | 18 | 4 | | |
| nb | Norwegian (Bokmål) | 641 | 140 | 0.84 | on UD NO-Bokmaal |
| nl | Dutch | 370 | 125 | 0.96 | on UD NL-Alpino, excl. underscore-joined compound lemmas |
| nn | Norwegian (Nynorsk) | 138 | 36 | 0.83 | on UD NO-Nynorsk |
| pl | Polish | 3,670 | 264 | 0.96 | on UD PL-LFG |
| pt | Portuguese | 927 | 95 | 0.95 | on UD PT-GSD |
| ro | Romanian | 345 | 37 | 0.94 | on UD RO-RRT |
| ru | Russian | 1,362 | 131 | 0.94 | on UD RU-SynTagRus, alternative: pymorphy2 |
| se | Northern Sámi | 115 | 7 | 0.97 | on UD SME-Giella |
| sk | Slovak | 908 | 73 | 0.92 | on UD SK-SNK |
| sl | Slovene | 157 | 31 | 0.95 | on UD SL-SSJ |
| sq | Albanian | 96 | 10 | 0.71 | on UD SQ-STAF |
| sv | Swedish | 964 | 129 | 0.95 | on UD SV-Talbanken, alternative: lemmy |
| sw | Swahili | 4,869 | 4 | | experimental |
| tl | Tagalog | 78 | 25 | 0.84 | on UD TL-TRG (test-only treebank, no train split) |
| tr | Turkish | 1,236 | 40 | 0.92 | on UD TR-KeNet |
| uk | Ukrainian | 599 | 45 | 0.94 | on UD UK-IU, alternative: pymorphy2 |
Languages marked as low-coverage may be better served by
language-specific libraries, which are referenced where an open-source
Python alternative exists. Simplemma still provides limited functionality.
Experimental means the language is untested, or that its data or
lemmatization may have issues.
The scores measure how accurately tokens are mapped to their lemma on
Universal Dependencies treebanks, over
single word tokens (including some contractions but not merged prepositions).
Each figure is the accuracy on the held-out dev+test splits of each language's
best-performing general-purpose treebank; train splits are excluded from
scoring as they are mined for the correction lists and gate every candidate.
The
training/ folder documents the protocol, the annotation-driven
exceptions (Dutch compound lemmas, Hebrew and Arabic proclitics,
Finnish/Estonian/Hungarian compound-boundary markers) and how to reproduce
the figures.
The benchmark only incidentally captures what this library is most useful
for, the lemmatization of less frequent words.
Roadmap
- [ ] Return all candidate lemmas for ambiguous words (#94, #132)
- [ ] Optional compound splitting (#141)
- [ ] More and better source data (#1, #3)
Credits and licenses
The software is licensed under the MIT license. For information on the
licenses of the linguistic information databases, see the
licenses` folder.
The surface lookups (non-greedy mode) rely on lemmatization lists derived from the following sources, listed in order of relative importance:
lists by Michal Měchura (Open Database License)- Wiktionary entries packaged by the Kaikki
Contributions
This package has been first created and published by Adrien Barbaresi. It has then benefited from extensive refactoring by Juanjo Diaz (especially the new classes). See the full list of contributors to the repository.
Feel free to contribute, notably by filing issues for feedback, bug reports, or links to further lemmatization lists, rules and tests.
Contributions by pull requests ought to follow the following conventions: code style and linting with ruff, type hinting with mypy, included tests with pytest.
Other solutions
See lists: German-NLP and other awesome-NLP lists.
For another approach in Python see spaCy's edit tree lemmatizer.
References
To cite this software:
Barbaresi A. (year). Simplemma: a simple multilingual lemmatizer for
Python [Computer software] (Version version number). Available from
This work draws from lexical analysis algorithms used in:
- Barbaresi, A., & Hein, K. (2017). Data-driven identification of
- Barbaresi, A. (2016). An unsupervised morphological criterion for
- Barbaresi, A. (2016). Bootstrapped OCR error detection for a