# How the Seedance 2.0 Multilingual Vocabulary System Works: A Deep Architectural Guide

> Explore the Seedance 2.0 multilingual vocabulary system architecture. Learn how plug-in skill files, load directives, and schema validation enable seamless language support.

- Repository: [Iamemily2050 /seedance-2.0](https://github.com/Emily2040/seedance-2.0)
- Tags: deep-dive
- Published: 2026-08-03

---

**Seedance 2.0 implements a plug-in-style multilingual vocabulary system where language-specific skill files declare term banks, the prompt compiler loads reference vocabularies via `load` directives, and schema validation ensures consistent structure across all supported languages.**

The [Emily2040/seedance-2.0](https://github.com/Emily2040/seedance-2.0) repository treats multilingual support as a modular skill layer rather than hardcoded translations. This design lets contributors add new languages without touching the core compiler logic.

## Core Components of the Multilingual Vocabulary System

### Language-Specific Skill Files ([`SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/SKILL.md))

Each language is encapsulated in its own skill directory under `skills/seedance-vocab-{xx}/`. The [`SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/SKILL.md) file serves as the contract between the language pack and the compiler.

**In [`skills/seedance-vocab-zh/SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/skills/seedance-vocab-zh/SKILL.md)**, the skill declares:

- **Metadata**: skill ID, version, tags (`chinese`, `vocabulary`)
- **Intent**: what the skill provides (Chinese cinematic terminology)
- **Term tables**: structured mappings for camera, lighting, motion, audio, and atmosphere
- **Compact pattern**: a dense phrasing template for efficient token usage
- **De-slop rules**: guardrails against vague or overused terms
- **Usage rules**: constraints like preserving reference tags (`@Image1`) unchanged

The compiler resolves these skills at runtime based on user requests.

### Reference Vocab Files (`references/vocab/*.md`)

Raw term data lives in portable Markdown files under `references/vocab/`. These files contain:

- Tabular vocabulary mappings (Chinese term → English gloss → usage notes)
- Pattern definitions for compact phrasing
- Script-variant rules for regional differences
- Slop-trap tables listing terms to avoid

Each vocab file is loaded via the `[ref:vocab/{code}]` directive in the skill's `Load` section. This decouples data from logic—skills define *how* to use terms, while reference files provide the *what*.

| File | Purpose |
|------|---------|
| [`references/vocab/zh.md`](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/zh.md) | Chinese (Simplified) cinematic vocabulary with regional variants |
| [`references/vocab/en.md`](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/en.md) | English base vocabulary used as fallback |
| [`references/vocab/ko.md`](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/ko.md) | Korean terminology and pattern rules |

### Schema Validation ([`scripts/vocab_schema_check.py`](https://github.com/Emily2040/seedance-2.0/blob/main/scripts/vocab_schema_check.py))

The [[`vocab_schema_check.py`](https://github.com/Emily2040/seedance-2.0/blob/main/vocab_schema_check.py)](https://github.com/Emily2040/seedance-2.0/blob/main/scripts/vocab_schema_check.py) script enforces structural consistency across all language files. It validates:

- Required table headers and column counts
- Presence of compact patterns for each category
- Valid slop-trap entries with suggested replacements
- Proper nesting of script variants

This prevents drift as the vocabulary system scales to additional languages.

## How the Prompt Compiler Integrates Multilingual Vocabularies

The compilation pipeline follows a strict resolution order defined in [[`V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md`](https://github.com/Emily2040/seedance-2.0/blob/main/V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md)](https://github.com/Emily2040/seedance-2.0/blob/main/V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md):

1. **Skill resolution** — Parse user request for language hints (explicit `skill:` declaration or inferred from input text)
2. **Vocab loading** — Execute `load: [ref:vocab/{xx}]` to pull term tables into the pipeline
3. **Rule application** — Enforce usage rules (preserve tags, apply de-slop constraints)
4. **Pattern compaction** — Inject compact phrasing templates appropriate to the language
5. **Final generation** — Emit the completed prompt with language-specific terminology

### Example Compilation Flow

```yaml
---

# User requests Chinese cinematic prompt

ask: "用中文写一个灯光设计，参考@Image1，保持主体不变，加入动态光效"

# Compiler resolves:

skill: seedance-vocab-zh
load: [ref:vocab/zh]
---

# Resolved internal representation:

@Image1 为绝对参考基准，[主体]严格保持；仅调整[灯光]。
布光：动态光效，强调轮廓分离。
镜头：固定机位，观察焦点。
氛围：戏剧性明暗对比。
声音：环境呼吸感。

```

The Chinese terms (`布光`, `镜头`, `氛围`) are drawn from the loaded `[ref:vocab/zh]` tables, while the structure follows the compact pattern defined in [`seedance-vocab-zh/SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/seedance-vocab-zh/SKILL.md).

## Adding a New Language: Step-by-Step

The modular architecture enables straightforward language expansion:

1. **Create skill directory**: `mkdir skills/seedance-vocab-{code}`
2. **Write [`SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/SKILL.md)**: Define intent, tag with appropriate language code, author compact pattern and rules
3. **Create reference file**: `references/vocab/{code}.md` with fully populated term tables
4. **Validate**: Run `python scripts/vocab_schema_check.py` to verify against schema
5. **Register**: Add language code to compiler's recognized skill index

No changes to the core compiler are required—the system discovers and loads new skills automatically.

## Practical Code Examples

### Invoking a Multilingual Vocabulary Skill

```yaml
---
skill: seedance-vocab-zh
load: [ref:vocab/zh]
---

@Image1 为参考基准，严格保持[主体]不变；仅加入[动作]。
镜头：缓慢推镜，焦点锁定主体。
声音：安静环境声，突出动作细节。

# Compiler produces: Chinese prompt with precise cinematic terminology

```

### Korean Vocabulary Skill Usage

```yaml
---
skill: seedance-vocab-ko
load: [ref:vocab/ko]
---

@Image1은 기준이며, [주제]를 유지하십시오; [동작]만 추가합니다.
카메라: 천천히 줌인, 피사체 고정.
사운드: 조용한 주변음.

```

### Programmatic Skill Resolution

```python
from seedance.compiler import PromptCompiler

compiler = PromptCompiler()

# Explicit language selection

compiled_zh = compiler.compile(
    skill="seedance-vocab-zh",
    load="[ref:vocab/zh]",
    template="@Image1 为参考...",
    constraints={"preserve_tags": ["@Image1"]}
)

# Language inference from input text

compiled_auto = compiler.compile(
    user_input="Créez un éclairage dramatique en français",
    auto_detect=True  # Resolves to seedance-vocab-fr if available

)

```

## Key Files and Their Roles

| Path | Function | GitHub Link |
|------|----------|-------------|
| [`skills/seedance-vocab-zh/SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/skills/seedance-vocab-zh/SKILL.md) | Chinese vocab skill definition with tables, patterns, rules | [View file](https://github.com/Emily2040/seedance-2.0/blob/main/skills/seedance-vocab-zh/SKILL.md) |
| [`skills/seedance-vocab-en/SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/skills/seedance-vocab-en/SKILL.md) | English fallback vocab skill | [View file](https://github.com/Emily2040/seedance-2.0/blob/main/skills/seedance-vocab-en/SKILL.md) |
| [`references/vocab/zh.md`](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/zh.md) | Raw Chinese term tables and pattern data | [View file](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/zh.md) |
| [`references/vocab/en.md`](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/en.md) | English base vocabulary tables | [View file](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/en.md) |
| [`scripts/vocab_schema_check.py`](https://github.com/Emily2040/seedance-2.0/blob/main/scripts/vocab_schema_check.py) | JSON schema validator for all vocab files | [View file](https://github.com/Emily2040/seedance-2.0/blob/main/scripts/vocab_schema_check.py) |
| [`V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md`](https://github.com/Emily2040/seedance-2.0/blob/main/V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md) | Compiler manifest defining skill loading and vocab resolution | [View file](https://github.com/Emily2040/seedance-2.0/blob/main/V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md) |

## Summary

- **Modular skill architecture** isolates each language in a self-contained `seedance-vocab-{xx}` directory with its own [`SKILL.md`](https://github.com/Emily2040/seedance-2.0/blob/main/SKILL.md)
- **Reference file loading** via `[ref:vocab/{code}]` directives decouples terminology data from skill logic
- **Schema validation** through [`vocab_schema_check.py`](https://github.com/Emily2040/seedance-2.0/blob/main/vocab_schema_check.py) maintains structural consistency as the system scales
- **Compiler integration** follows a resolution pipeline: skill detection → vocab loading → rule application → pattern compaction → final generation
- **Zero core changes** required to add languages—simply create skill and reference files following the established patterns

## Frequently Asked Questions

### How does Seedance 2.0 detect which language to use?

The compiler checks for explicit `skill:` declarations first. If absent, it analyzes the input text for language-specific character ranges and keywords, then matches against available `seedance-vocab-{xx}` skills. The detection logic weights Unicode block presence (e.g., CJK for Chinese/Japanese/Korean, Hangul for Korean) against registered skill tags.

### What happens if a requested language has no vocabulary skill?

The compiler falls back to `seedance-vocab-en` (English) and emits a warning log. The [[`V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md`](https://github.com/Emily2040/seedance-2.0/blob/main/V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md)](https://github.com/Emily2040/seedance-2.0/blob/main/V6_SEQUENCE_PROMPT_COMPILER_MANIFEST.md) specifies that compilation must never fail due to missing localization—output degrades gracefully to the base language while preserving all structural constraints.

### Can vocabulary terms reference each other across languages?

No—cross-language term linking is intentionally prohibited. Each `[ref:vocab/{code}]` loads an isolated namespace. If a prompt requires mixed-language output (e.g., Chinese description with English technical tags), the skill's **Usage Rule** explicitly defines which elements remain untranslated, rather than pulling from multiple vocab files simultaneously.

### How are regional variants (Simplified vs. Traditional Chinese) handled?

Script variants are defined within a single vocab file using variant tables. In [[`references/vocab/zh.md`](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/zh.md)](https://github.com/Emily2040/seedance-2.0/blob/main/references/vocab/zh.md), terms include `variant-simplified` and `variant-traditional` columns. The skill's compact pattern selects the appropriate column based on a `script:` parameter in the user request or defaults to Simplified for mainland Chinese contexts.