# How VoiceStudio Dubbing Translator Maps ISO Codes to Engine-Specific Locales

> Learn how VoiceStudio's dubbing translator maps ISO codes to engine-specific locales. Achieve seamless interoperability across cloud APIs, NLLB, and LLM providers.

- Repository: [Palash Debnath/VoiceStudio](https://github.com/debpalash/VoiceStudio)
- Tags: internals
- Published: 2026-09-13

---

**VoiceStudio's dubbing translation endpoint normalizes ISO-639-1 language identifiers into engine-specific locale strings using specialized mapping dictionaries, enabling seamless interoperability between cloud APIs, offline NLLB models, and LLM providers.**

The VoiceStudio open-source repository implements a sophisticated locale resolution system within its dubbing pipeline to handle diverse translation backends. Located in [`backend/api/routers/dub_translate.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/dub_translate.py), the translation router converts standard two-letter language codes into the precise formats required by each engine—from Google's cloud API to Meta's No Language Left Behind (NLLB) model. This mapping ensures that source and target languages are correctly interpreted regardless of which provider handles the actual text translation.

## Cloud Provider Locale Resolution via TRANSLATE_CODES

For cloud-based and legacy translation engines including Google Translate, DeepL, and Argos-Translate, VoiceStudio uses the `TRANSLATE_CODES` dictionary to map ISO-639-1 codes to provider-expected locale strings.

The mapping table handles standard abbreviations and specific locale variants required by certain engines. For example, simplified Chinese `zh` maps to `zh-CN` for Argos compatibility, while most European languages pass through unchanged.

```python
TRANSLATE_CODES = {
    "en": "en", "es": "es", "fr": "fr", "de": "de", "it": "it", "pt": "pt",
    "ru": "ru", "ja": "ja", "ko": "ko", "zh": "zh-CN", "cmn-Hans": "zh-CN",
    "ar": "ar", "hi": "hi", "tr": "tr", "pl": "pl", "nl": "nl", "sv": "sv",
    "th": "th", "vi": "vi", "id": "id", "uk": "uk",
}

```

When processing a translation request, the handler looks up the client-supplied `target_lang` in this dictionary (lines 28-33 of [`dub_translate.py`](https://github.com/debpalash/VoiceStudio/blob/main/dub_translate.py)). The resolved `lang_code` value is then passed to the selected provider's translation method. For Argos-Translate, this ensures the engine receives `zh-CN` rather than the generic `zh` code, preventing locale mismatches that would cause translation failures.

## FLORES-200 Mapping for NLLB Offline Translation

When the `provider` parameter is set to `"nllb"`, VoiceStudio switches to the `FLORES_CODES` dictionary to support Meta's NLLB-200 model. This transformer architecture requires FLORES-200 language identifiers rather than standard ISO codes.

The mapping converts ISO-639-1 codes to FLORES format, which combines ISO-639-3 three-letter codes with script identifiers. For instance, `zh` becomes `zho_Hans` for simplified Chinese, and `ar` becomes `arb_Arab` for Arabic.

```python
FLORES_CODES = {
    "en": "eng_Latn", "es": "spa_Latn", "fr": "fra_Latn", "de": "deu_Latn",
    "it": "ita_Latn", "pt": "por_Latn", "ru": "rus_Cyrl", "ja": "jpn_Jpan",
    "ko": "kor_Hang", "zh": "zho_Hans", "zh-CN": "zho_Hans",
    "cmn-Hans": "zho_Hans", "ar": "arb_Arab", "hi": "hin_Deva",
    "tr": "tur_Latn", "pl": "pol_Latn", "nl": "nld_Latn",
    "sv": "swe_Latn", "th": "tha_Thai", "vi": "vie_Latn",
    "id": "ind_Latn", "uk": "ukr_Cyrl",
}

```

During NLLB inference (lines 84-86), the system retrieves both source and target language codes from this mapping. The `flores_src` and `flores_tgt` variables are then passed to the transformer model to ensure proper tokenization and translation directionality.

## LLM Provider Integration with Natural Language Names

For LLM-based providers such as OpenAI, VoiceStudio bypasses the locale code translation entirely. Instead, the system uses the `LANG_NAMES` dictionary to inject natural language descriptions directly into the system prompt.

Unlike traditional translation engines that require specific locale strings, LLMs understand natural language instructions. The router passes the raw ISO code to generate phrases like "translate to Spanish" rather than manipulating machine-readable locale identifiers. This approach allows the LLM to handle nuanced instructions without strict formatting constraints.

## Dialect-Aware Translation Without Locale Mutation

VoiceStudio supports dialect-specific translations through the `DIALECT_HINTS` table, which operates independently of the locale mapping system. When a request includes a `dialect` parameter (e.g., `es-AR` for Argentinian Spanish), the system appends a dialect clause to the LLM prompt.

This mechanism does not alter the underlying locale identifier passed to the translation engine. For cloud providers, the base language code (e.g., `es`) remains unchanged; instead, the dialect hint influences tone, vocabulary, and regional variations through prompt engineering. The `dialect_clause` is constructed from `DIALECT_HINTS` and prepended to the translation context.

## Engine Selection and Routing Logic

The translation router determines which mapping strategy to apply based on the `req.provider` field, defaulting to `"google"` if unspecified. The selection logic follows this priority:

1. **NLLB provider**: Uses `FLORES_CODES` for offline transformer inference
2. **Argos or legacy providers**: Uses `TRANSLATE_CODES` for standardized locale conversion
3. **OpenAI/LLM providers**: Uses raw ISO codes with `LANG_NAMES` for natural language prompting

This architecture allows the same `TranslateRequest` schema defined in [`schemas/requests.py`](https://github.com/debpalash/VoiceStudio/blob/main/schemas/requests.py) to function across heterogeneous backends without requiring clients to understand each engine's specific locale requirements.

## Practical Implementation Examples

The following examples demonstrate how VoiceStudio handles different translation scenarios:

**Argos-Translate with Chinese Simplification:**

```python
req = TranslateRequest(
    provider="argos",
    target_lang="zh",          # ISO code supplied by the client

    segments=[Segment(id=1, text="Hello world!")],
)

# Internally:

lang_code = TRANSLATE_CODES.get(req.target_lang, req.target_lang)  # → "zh-CN"

# Argos-Translate receives from_code="en" and to_code="zh-CN"

```

**NLLB Offline Model:**

```python
req.provider = "nllb"
src_lang = "en"
target_lang = "zh"
flores_src = FLORES_CODES.get(src_lang, "eng_Latn")   # "eng_Latn"

flores_tgt = FLORES_CODES.get(target_lang, "eng_Latn")  # "zho_Hans"

# NLLB model is invoked with src_lang and tgt_lang set to FLORES identifiers

```

**OpenAI with Dialect Specification:**

```python
req = TranslateRequest(
    provider="openai",
    target_lang="es",
    dialect="es-AR",
    segments=[Segment(id=1, text="How are you?")],
)

# The system prompt will include:

#   "Target dialect — Rioplatense Spanish (Argentina): use voseo ..."

# but the locale identifier stays "es".

```

## Summary

- VoiceStudio maintains two primary mapping dictionaries in [`backend/api/routers/dub_translate.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/dub_translate.py): `TRANSLATE_CODES` for cloud providers and `FLORES_CODES` for NLLB models.
- Cloud translation engines (Google, DeepL, Argos) receive locale strings like `zh-CN` through the `TRANSLATE_CODES` lookup table.
- NLLB offline inference requires FLORES-200 codes (e.g., `zho_Hans`, `eng_Latn`) resolved via the `FLORES_CODES` dictionary at lines 84-86.
- LLM providers use natural language names derived from ISO codes rather than technical locale identifiers.
- Dialect hints modify LLM prompts without changing the underlying engine locale, enabling regional variations like Rioplatense Spanish.

## Frequently Asked Questions

### How does VoiceStudio handle unsupported ISO codes?

When an ISO code is not present in `TRANSLATE_CODES` or `FLORES_CODES`, the system falls back to using the raw client-supplied value via `.get(code, code)` semantics. This allows custom or newly added languages to pass through to the provider, though success depends on the specific engine's support for that identifier.

### What is the difference between TRANSLATE_CODES and FLORES_CODES?

`TRANSLATE_CODES` maps ISO-639-1 codes to provider-specific locale strings used by commercial APIs and Argos-Translate, such as converting `zh` to `zh-CN`. `FLORES_CODES` converts the same ISO codes to FLORES-200 format (e.g., `zh` to `zho_Hans`) required by Meta's NLLB-200 transformer architecture for offline inference.

### Where is the translation engine routing logic implemented?

The routing logic resides in [`backend/api/routers/dub_translate.py`](https://github.com/debpalash/VoiceStudio/blob/main/backend/api/routers/dub_translate.py), specifically between lines 28-42 and 84-86. The `req.provider` field determines whether the system consults `TRANSLATE_CODES`, `FLORES_CODES`, or passes raw ISO codes to LLM prompts. Supporting utilities for engine management are located in [`services/translation_engines.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/translation_engines.py) and [`services/model_manager.py`](https://github.com/debpalash/VoiceStudio/blob/main/services/model_manager.py).

### Does VoiceStudio support regional dialects in machine translation?

Yes, VoiceStudio supports dialects through the `DIALECT_HINTS` system when using LLM providers like OpenAI. The dialect parameter (e.g., `es-AR` for Argentinian Spanish) injects specific instructions into the system prompt to influence regional vocabulary and grammar, while the base ISO code remains unchanged for the translation engine itself.