Supported Languages in PaddleOCR: A Complete Guide to 100+ Multilingual Models
PaddleOCR supports over 100 languages across its model families, including 80 languages in the base multilingual model, 106 in PP-OCRv5, and 109 in PaddleOCR-VL, configurable via the lang parameter in Python or --lang flag in CLI.
The PaddlePaddle/PaddleOCR repository delivers production-grade optical character recognition for diverse global scripts. Understanding the supported languages in PaddleOCR is essential for accurate text recognition, as the framework handles Latin, Cyrillic, Asian, and right-to-left languages through specific abbreviation codes defined in the source documentation and character dictionaries.
Language Coverage by Model Generation
PaddleOCR organizes language support across multiple model families, with each generation expanding the number of available scripts.
PP-OCRv3 and PP-OCRv4 (80 Languages)
The core multilingual recognition models support 80 languages as documented in docs/version2.x/ppocr/blog/multi_languages.en.md. This baseline set covers major world languages including Chinese, English, Japanese, Korean, Arabic, and numerous European and Indic scripts.
PP-OCRv5 (106 Languages)
The latest PP-OCRv5 generation expands coverage to 106 languages, adding support for additional Latin-script languages and low-resource dialects. The complete enumeration resides in docs/version3.x/algorithm/PP-OCRv5/PP-OCRv5_multi_languages.en.md.
PaddleOCR-VL (109 Languages)
The vision-language variant, PaddleOCR-VL, provides the broadest coverage with 109 languages, integrating multilingual capabilities with visual understanding as detailed in docs/version3.x/algorithm/PaddleOCR-VL/PaddleOCR-VL.en.md.
Language Abbreviation Reference Table
The following table presents the complete mapping of language names to their corresponding abbreviation codes used in the PaddleOCR API. These abbreviations are passed to the lang parameter in Python or the --lang flag in CLI commands.
| Language | Abbreviation | Language | Abbreviation |
|---|---|---|---|
| Chinese & English | ch |
Arabic | ar |
| English | en |
Hindi | hi |
| French | fr |
Uyghur | ug |
| German | german |
Persian | fa |
| Japanese | japan |
Urdu | ur |
| Korean | korean |
Serbian (latin) | rs_latin |
| Chinese Traditional | chinese_cht |
Occitan | oc |
| Italian | it |
Marathi | mr |
| Spanish | es |
Nepali | ne |
| Portuguese | pt |
Serbian (cyrillic) | rs_cyrillic |
| Russian | ru |
Bulgarian | bg |
| Ukrainian | uk |
Estonian | et |
| Belarusian | be |
Irish | ga |
| Telugu | te |
Croatian | hr |
| Sanskrit | sa |
Hungarian | hu |
| Tamil | ta |
Indonesian | id |
| Afrikaans | af |
Icelandic | is |
| Azerbaijani | az |
Kurdish | ku |
| Bosnian | bs |
Lithuanian | lt |
| Czech | cs |
Latvian | lv |
| Welsh | cy |
Maori | mi |
| Danish | da |
Malay | ms |
| Maltese | mt |
Adyghe | ady |
| Dutch | nl |
Kabardian | kbd |
| Norwegian | no |
Avar | ava |
| Polish | pl |
Dargwa | dar |
| Romanian | ro |
Ingush | inh |
| Slovak | sk |
Lak | lbe |
| Slovenian | sl |
Lezghian | lez |
| Albanian | sq |
Tabassaran | tab |
| Swedish | sv |
Bihari | bh |
| Swahili | sw |
Maithili | mai |
| Tagalog | tl |
Angika | ang |
| Turkish | tr |
Bhojpuri | bho |
| Uzbek | uz |
Magahi | mah |
| Vietnamese | vi |
Nagpur | sck |
| Mongolian | mn |
Newari | new |
| Abaza | abq |
Goan Konkani | `gom |
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →