# How Per-Stage Model Routing Selects Different Models in OpenMAIC

> Discover how per-stage model routing in OpenMAIC selects specific LLMs for each generation phase using a JSON config. Learn about runtime resolution and default model fallback.

- Repository: [MAIC/OpenMAIC](https://github.com/THU-MAIC/OpenMAIC)
- Tags: internals
- Published: 2026-09-13

---

**Per-stage model routing in OpenMAIC maps generation phases (system, assistant, tool, final) to specific LLMs via a JSON configuration, resolving the appropriate model at runtime through [`lib/server/resolve-model.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/resolve-model.ts) while falling back to a default model when no stage-specific mapping exists.**

OpenMAIC supports **per-stage model routing**, an architecture that lets you assign distinct language models to individual phases of a generation workflow. This mechanism reads a JSON routing table from the `MODEL_ROUTES` environment variable at startup, validates the configuration, and dynamically selects the optimal LLM for each generation request. By leveraging this system, you can route system prompts to instruction-tuned models and tool calls to specialized reasoning models within a single interaction.

## Configuring Stage-to-Model Mappings

The routing configuration is defined as a JSON object assigned to the `MODEL_ROUTES` environment variable, typically set in your `.env` file. This object maps **generation stages** to **model identifiers** or provider aliases.

The supported stage names are:
- **system** – Initial instruction and context setting
- **assistant** – Primary response generation and code synthesis  
- **tool** – Tool use and function calling reasoning
- **final** – Final answer synthesis and output formatting

Example configuration from `.env.example`:

```dotenv
MODEL_ROUTES={"system":"glm-5.2","assistant":"kimi-k2.7","tool":"qwen3.7-max"}

```

In this setup, GLM-5.2 handles system prompts for strong instruction adherence, Kimi-K2.7 generates assistant responses optimized for code, and Qwen-3.7-Max processes tool calls.

## Boot-Time Validation of Routing Tables

Before the server accepts traffic, [`lib/server/config-validation.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/config-validation.ts) performs a **warn-first** validation of the `MODEL_ROUTES` configuration. This validation step checks for unknown stage names and verifies that specified models are available on the configured provider. Warnings are emitted for misconfigurations, preventing invalid routing tables from crashing the server while alerting operators to potential issues.

The validation logic ensures that:
- All keys in the JSON object correspond to valid generation stages
- Referenced models exist in the provider's available model list
- The JSON structure is parseable and type-safe

## Runtime Model Resolution

When a generation request enters the system, the core resolver in [`lib/server/resolve-model.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/resolve-model.ts) determines which LLM to invoke. The resolver receives the current **generation stage** from the `@openmaic/generation` package and performs a lookup against the routing table defined in [`lib/server/model-routes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/model-routes.ts).

### Stage Lookup and Default Fallback

The resolution logic follows a strict priority:
1. If the stage exists in `MODEL_ROUTES`, return the configured model identifier
2. If no mapping exists, fall back to the **default model** defined in the global configuration
3. If the stage is explicitly `null` or undefined, bypass routing entirely and use the default model

This **per-request** resolution allows mixed-model interactions without server restarts. For example, a single conversation can use GLM-5.2 for the system phase, then switch to Kimi-K2.7 for the assistant phase based on the stage parameter passed to the resolver.

### Cross-Provider Model Translation

The resolver supports **cross-provider routing** by translating generic model aliases into provider-specific identifiers. For instance, the alias `"glm-5.2"` can resolve to OpenAI's `gpt-4` or another provider's equivalent, depending on the active provider configuration. The test suite in [`tests/server/resolve-model.test.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/tests/server/resolve-model.test.ts) validates both stage-specific provider selection and the fallback behavior when no stage is supplied.

## Integration with the Generation Pipeline

The routing mechanism integrates with the generation library through a **neutral callback seam** implemented in [`generation-ai-call.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/generation-ai-call.ts). When the `@openmaic/generation` package initiates a call, it emits the current stage name through this seam. The server-side routing logic then applies uniformly across all generation types, including planner steps, tool invocations, and streaming responses.

This architecture decouples the generation logic from model selection, allowing the routing layer to evolve independently of the core generation algorithms.

## Implementation Examples

### Accessing the Routing Table in Code

Use the `getModelForStage` helper from `@openmaic/server` to resolve models programmatically:

```typescript
import { getModelForStage } from '@openmaic/server';

async function generate(stage: GenerationStage, prompt: string) {
  const model = getModelForStage(stage); // Resolves based on MODEL_ROUTES
  return await callModel(model, prompt);
}

```

### Runtime Override for Specific Sessions

You can override routing dynamically for individual sessions without modifying the global environment:

```typescript
const sessionRoutes = { ...process.env.MODEL_ROUTES, assistant: 'gpt-4o' };
setModelRoutes(sessionRoutes);

```

This pattern enables A/B testing of models or user-specific routing preferences while maintaining the default configuration for other requests.

## Summary

- **Per-stage model routing** routes distinct generation phases (system, assistant, tool, final) to specialized LLMs via the `MODEL_ROUTES` environment variable.
- The [`lib/server/config-validation.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/config-validation.ts) module performs boot-time validation to catch configuration errors before they affect traffic.
- [`lib/server/resolve-model.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/resolve-model.ts) provides the runtime resolution logic, falling back to the default model when a stage lacks a specific mapping.
- The **neutral callback seam** in [`generation-ai-call.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/generation-ai-call.ts) ensures uniform routing across all generation calls from the `@openmaic/generation` package.
- Routing decisions are made **per request**, enabling dynamic model selection without server restarts.

## Frequently Asked Questions

### What happens if a generation stage is not defined in MODEL_ROUTES?

If a stage lacks a specific mapping in the `MODEL_ROUTES` configuration, the resolver in [`lib/server/resolve-model.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/resolve-model.ts) automatically falls back to the **default model** defined in the global server configuration. This ensures that generation requests always proceed, even with incomplete routing tables.

### Can I route different stages to models from different AI providers?

Yes. The resolver supports **cross-provider routing** by translating model aliases into provider-specific identifiers. You can configure the system stage to use a model from Provider A and the tool stage to use a model from Provider B, with the resolver handling the provider switch transparently at runtime.

### Is the routing configuration hot-reloadable without restarting the server?

While the base `MODEL_ROUTES` environment variable is read at startup, you can override routing at runtime using the `setModelRoutes` function. This allows session-specific or dynamic routing changes without requiring a server restart, though persistent changes should be updated in the environment configuration.

### How does per-stage routing affect request latency?

The routing logic adds negligible overhead because resolution occurs through a simple object lookup in [`lib/server/model-routes.ts`](https://github.com/THU-MAIC/OpenMAIC/blob/main/lib/server/model-routes.ts). The **per-request** decision happens once per generation phase, and the system does not introduce additional network calls or heavy computation during the routing phase.