How Per-Stage Model Routing Selects Different Models in OpenMAIC

Per-stage model routing in OpenMAIC maps generation phases (system, assistant, tool, final) to specific LLMs via a JSON configuration, resolving the appropriate model at runtime through lib/server/resolve-model.ts while falling back to a default model when no stage-specific mapping exists.

OpenMAIC supports per-stage model routing, an architecture that lets you assign distinct language models to individual phases of a generation workflow. This mechanism reads a JSON routing table from the MODEL_ROUTES environment variable at startup, validates the configuration, and dynamically selects the optimal LLM for each generation request. By leveraging this system, you can route system prompts to instruction-tuned models and tool calls to specialized reasoning models within a single interaction.

Configuring Stage-to-Model Mappings

The routing configuration is defined as a JSON object assigned to the MODEL_ROUTES environment variable, typically set in your .env file. This object maps generation stages to model identifiers or provider aliases.

The supported stage names are:

  • system – Initial instruction and context setting
  • assistant – Primary response generation and code synthesis
  • tool – Tool use and function calling reasoning
  • final – Final answer synthesis and output formatting

Example configuration from .env.example:

MODEL_ROUTES={"system":"glm-5.2","assistant":"kimi-k2.7","tool":"qwen3.7-max"}

In this setup, GLM-5.2 handles system prompts for strong instruction adherence, Kimi-K2.7 generates assistant responses optimized for code, and Qwen-3.7-Max processes tool calls.

Boot-Time Validation of Routing Tables

Before the server accepts traffic, lib/server/config-validation.ts performs a warn-first validation of the MODEL_ROUTES configuration. This validation step checks for unknown stage names and verifies that specified models are available on the configured provider. Warnings are emitted for misconfigurations, preventing invalid routing tables from crashing the server while alerting operators to potential issues.

The validation logic ensures that:

  • All keys in the JSON object correspond to valid generation stages
  • Referenced models exist in the provider's available model list
  • The JSON structure is parseable and type-safe

Runtime Model Resolution

When a generation request enters the system, the core resolver in lib/server/resolve-model.ts determines which LLM to invoke. The resolver receives the current generation stage from the @openmaic/generation package and performs a lookup against the routing table defined in lib/server/model-routes.ts.

Stage Lookup and Default Fallback

The resolution logic follows a strict priority:

  1. If the stage exists in MODEL_ROUTES, return the configured model identifier
  2. If no mapping exists, fall back to the default model defined in the global configuration
  3. If the stage is explicitly null or undefined, bypass routing entirely and use the default model

This per-request resolution allows mixed-model interactions without server restarts. For example, a single conversation can use GLM-5.2 for the system phase, then switch to Kimi-K2.7 for the assistant phase based on the stage parameter passed to the resolver.

Cross-Provider Model Translation

The resolver supports cross-provider routing by translating generic model aliases into provider-specific identifiers. For instance, the alias "glm-5.2" can resolve to OpenAI's gpt-4 or another provider's equivalent, depending on the active provider configuration. The test suite in tests/server/resolve-model.test.ts validates both stage-specific provider selection and the fallback behavior when no stage is supplied.

Integration with the Generation Pipeline

The routing mechanism integrates with the generation library through a neutral callback seam implemented in generation-ai-call.ts. When the @openmaic/generation package initiates a call, it emits the current stage name through this seam. The server-side routing logic then applies uniformly across all generation types, including planner steps, tool invocations, and streaming responses.

This architecture decouples the generation logic from model selection, allowing the routing layer to evolve independently of the core generation algorithms.

Implementation Examples

Accessing the Routing Table in Code

Use the getModelForStage helper from @openmaic/server to resolve models programmatically:

import { getModelForStage } from '@openmaic/server';

async function generate(stage: GenerationStage, prompt: string) {
  const model = getModelForStage(stage); // Resolves based on MODEL_ROUTES
  return await callModel(model, prompt);
}

Runtime Override for Specific Sessions

You can override routing dynamically for individual sessions without modifying the global environment:

const sessionRoutes = { ...process.env.MODEL_ROUTES, assistant: 'gpt-4o' };
setModelRoutes(sessionRoutes);

This pattern enables A/B testing of models or user-specific routing preferences while maintaining the default configuration for other requests.

Summary

  • Per-stage model routing routes distinct generation phases (system, assistant, tool, final) to specialized LLMs via the MODEL_ROUTES environment variable.
  • The lib/server/config-validation.ts module performs boot-time validation to catch configuration errors before they affect traffic.
  • lib/server/resolve-model.ts provides the runtime resolution logic, falling back to the default model when a stage lacks a specific mapping.
  • The neutral callback seam in generation-ai-call.ts ensures uniform routing across all generation calls from the @openmaic/generation package.
  • Routing decisions are made per request, enabling dynamic model selection without server restarts.

Frequently Asked Questions

What happens if a generation stage is not defined in MODEL_ROUTES?

If a stage lacks a specific mapping in the MODEL_ROUTES configuration, the resolver in lib/server/resolve-model.ts automatically falls back to the default model defined in the global server configuration. This ensures that generation requests always proceed, even with incomplete routing tables.

Can I route different stages to models from different AI providers?

Yes. The resolver supports cross-provider routing by translating model aliases into provider-specific identifiers. You can configure the system stage to use a model from Provider A and the tool stage to use a model from Provider B, with the resolver handling the provider switch transparently at runtime.

Is the routing configuration hot-reloadable without restarting the server?

While the base MODEL_ROUTES environment variable is read at startup, you can override routing at runtime using the setModelRoutes function. This allows session-specific or dynamic routing changes without requiring a server restart, though persistent changes should be updated in the environment configuration.

How does per-stage routing affect request latency?

The routing logic adds negligible overhead because resolution occurs through a simple object lookup in lib/server/model-routes.ts. The per-request decision happens once per generation phase, and the system does not introduce additional network calls or heavy computation during the routing phase.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →