How OpenClaude Derives Runtime Metadata for LLM Requests: Environment-to-Request Pipeline

OpenClaude derives runtime metadata for LLM requests through a deterministic four-step pipeline that aggregates environment variables, resolves routing configurations, merges catalog entries, and infers provider-specific OpenAI-shim settings to produce a complete OpenAIShimRuntimeContext object.

OpenClaude implements a sophisticated configuration resolution system to standardize interactions with diverse LLM providers. When preparing an API call, the framework must derive comprehensive runtime metadata for LLM requests that captures routing logic, model capabilities, and transport-specific overrides. This process centers on src/integrations/runtimeMetadata.ts, which orchestrates environment detection, route resolution, and configuration merging to ensure every request carries the correct parameters.

The Four-Step Derivation Pipeline

The runtime metadata derivation follows a strict logical sequence defined in src/integrations/runtimeMetadata.ts, transforming raw environment inputs into structured execution context.

Step 1: Gathering Environment Information

The pipeline begins with resolveOpenAIShimRuntimeContext, which reads the current process environment (or an optional processEnv override) to capture critical variables such as OPENAI_BASE_URL and OPENAI_MODEL. This function establishes the baseline configuration by detecting any explicit user overrides before proceeding to route resolution.

Step 2: Resolving the Active Route

Once environment variables are captured, the system determines which provider route to use through resolveActiveRouteIdFromEnv. This function inspects OPENAI_ROUTE_ID, OPENAI_PROFILE, and provider-specific environment variables to locate the appropriate route.

If a baseUrl parameter is supplied, resolveRouteIdFromBaseUrl can force route selection to match that endpoint, while preferBaseUrlRoute implements precedence logic that prioritizes explicit base URLs over profile-based configurations.

Step 3: Pulling Catalog and Descriptor Data

With a finalized routeId, the resolver pulls three critical data structures from the integration registry:

  • RouteDescriptor via getRouteDescriptor: Contains static transport configuration including default OpenAI-shim settings.
  • ModelCatalogEntry via getCatalogEntryForModel: Matches the model name against the catalog, resolving aliases and provider-mapped identifiers.
  • ModelDescriptor via getModelDescriptorForCatalogEntry: Provides provider-specific context window sizes and maximum output token limits.

Step 4: Assembling the Final Runtime Metadata

The final assembly phase merges three layers of OpenAI-shim configuration using mergeOpenAIShimConfig:

  1. Base configuration from the route descriptor (descriptor?.transportConfig.openaiShim).
  2. Catalog overrides from the entry (catalogEntry?.transportOverrides?.openaiShim).
  3. Inferred configuration via inferRemoteModelOpenAIShimConfig, which applies heuristics based on model strings (e.g., glm-5.2, deepseek, kimi) to set fields like preserveReasoningContent, maxTokensField, and removeBodyFields.

After merging, resolveRouteOpenAIShimConfig sanitizes attribution headers (such as those for the aimlapi partner) and returns a complete OpenAIShimRuntimeContext object containing the resolved routeId, descriptor, catalogEntry, and the fully merged openaiShimConfig.

From Metadata to Request Limits

The derived runtime metadata feeds directly into model limit resolution through resolveModelRuntimeLimits in src/integrations/runtimeMetadata.ts. This function extracts concrete contextWindow and maxOutputTokens values from the ModelDescriptor, falling back to the catalog entry limits if the descriptor lacks specifications, and finally defaulting to global constants like CLAUDE_CODE_MAX_OUTPUT_TOKENS for Claude models or permissive upper bounds for unknown providers.

In src/utils/context.ts, the buildRequestContext function orchestrates this integration:

// src/utils/context.ts (excerpt)
import { resolveOpenAIShimRuntimeContext } from '../integrations/runtimeMetadata.js';
import { resolveModelRuntimeLimits } from '../integrations/runtimeMetadata.js';

export async function buildRequestContext(opts) {
  const shimCtx = resolveOpenAIShimRuntimeContext(opts);
  const limits   = await resolveModelRuntimeLimits(shimCtx.catalogEntry);
  // …use shimCtx.openaiShimConfig + limits to shape the payload
}

Practical Implementation Examples

Manual Resolution for a Gemini Model

// Example 1 – Manually resolve a runtime context for a Gemini model
import { resolveOpenAIShimRuntimeContext } from '@/integrations/runtimeMetadata.js';

const ctx = resolveOpenAIShimRuntimeContext({
  baseUrl: 'https://generativelanguage.googleapis.com',
  model:   'gemini-1.5-pro',
  treatAsLocal: false,
});

console.log(ctx.routeId);               // “gemini”
console.log(ctx.catalogEntry?.id);       // “gemini-1.5-pro”
console.log(ctx.openaiShimConfig);      // merged shim config for Gemini

Using Context During Request Preparation

// Example 2 – Using the context while sending a request
import { buildRequestContext } from '@/utils/context.js';
import { requestExecutor } from '@/services/api/openaiShim/requestExecutor.js';

const requestOpts = {
  baseUrl: process.env.OPENAI_BASE_URL,
  model:   process.env.OPENAI_MODEL,
};

const ctx = await buildRequestContext(requestOpts);
const response = await requestExecutor({
  url: `${ctx.openaiShimConfig.baseUrl}/v1/chat/completions`,
  body: {
    model: ctx.openaiShimConfig.model ?? requestOpts.model,
    max_tokens: ctx.openaiShimConfig.maxTokensField
      ? ctx.openaiShimConfig.maxTokensField
      : undefined,
    // …other fields are taken from ctx.openaiShimConfig
  },
});

Inspecting Derived Limits Directly

// Example 3 – Inspecting the derived limits directly
import { resolveModelRuntimeLimits } from '@/integrations/runtimeMetadata.js';

const limits = await resolveModelRuntimeLimits(
  // Provide a catalog entry (or null) to get concrete limits
  {
    modelDescriptorId: 'gemini-1.5-pro',
    // …
  },
);
console.log(limits.contextWindow);   // e.g. 8192
console.log(limits.maxOutputTokens); // e.g. 2048

Summary

  • OpenClaude derives runtime metadata for LLM requests through a deterministic pipeline defined in src/integrations/runtimeMetadata.ts.
  • The process follows four steps: environment gathering, route resolution, catalog data retrieval, and multi-layer configuration merging.
  • Route descriptors provide base transport settings, while model catalog entries supply overrides and capability limits.
  • The inferRemoteModelOpenAIShimConfig heuristic automatically configures provider-specific fields based on model name patterns.
  • Final metadata consumption occurs in src/utils/context.ts, which combines shim configuration with resolved model limits to shape API payloads.
  • Unit tests in src/integrations/runtimeMetadata.test.ts and src/integrations/runtimeMetadata.modelLimits.test.ts verify correct merging, inference, and fallback behaviors.

Frequently Asked Questions

What is the entry point for deriving runtime metadata in OpenClaude?

The primary entry point is resolveOpenAIShimRuntimeContext in src/integrations/runtimeMetadata.ts. This function accepts an optional environment override and returns a complete OpenAIShimRuntimeContext object containing the resolved route, catalog entry, and merged OpenAI-shim configuration.

How does OpenClaude handle conflicting route configurations between environment variables and base URLs?

When both a profile (via OPENAI_PROFILE) and a baseUrl are present, the preferBaseUrlRoute logic prioritizes the explicit base URL over profile-based routing. The resolveRouteIdFromBaseUrl function can force route selection to match the provided endpoint URL, ensuring direct API endpoints take precedence over configured profiles.

Where are model-specific token limits defined in the runtime metadata pipeline?

Token limits originate from the ModelDescriptor retrieved via getModelDescriptorForCatalogEntry in src/integrations/runtimeMetadata.ts. If the descriptor lacks explicit values, resolveModelRuntimeLimits falls back to the ModelCatalogEntry specifications, then to global defaults like CLAUDE_CODE_MAX_OUTPUT_TOKENS for Claude models or permissive upper bounds for unknown providers.

How does the system infer configuration for unknown or remote models?

The inferRemoteModelOpenAIShimConfig function in src/integrations/runtimeMetadata.ts analyzes model name strings (such as deepseek, kimi, or moonshot) to automatically apply provider-specific settings. This heuristic determines fields like preserveReasoningContent, maxTokensField, and removeBodyFields without requiring explicit catalog entries for every possible model variant.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →