How Switchyard Exposes Codex-Compatible Model Metadata on the /v1/models Endpoint
Switchyard returns a dual-format JSON payload from GET /v1/models that contains both a generic OpenAI-style model catalog in the data field and a Codex-specific model list in the models field, sourced from the router's runtime configuration.
The NVIDIA-NeMo/Switchyard repository implements this endpoint in the native HTTP server (the switchyard-server crate) to advertise available AI models while satisfying the distinct schema requirements of Codex clients. The response merges standard API compatibility with Codex-specific metadata fields, enabling the ModelInfo cards that Codex-based tools expect.
The Dual-Format Response Structure
Switchyard's /v1/models endpoint returns a single JSON object containing two parallel representations of the available model fleet:
data– A generic list built bymodel_entry_json()that follows the OpenAIlistobject format, containing standard fields likeid,object, andcapabilities.models– A Codex-specific list built bycodex_model_entry_json()that includes extended metadata required by Codex clients, such asslug,shell_type,default_reasoning_level, andtruncation_policy.
This structure allows standard HTTP clients to consume the familiar OpenAI-style catalog while Codex-specific tooling can access the enriched metadata needed for shell command generation and reasoning configuration.
Endpoint Implementation in switchyard-server
The handler for GET /v1/models is defined in crates/switchyard-server/src/lib.rs at lines 60–66:
// GET /v1/models → returns a JSON payload
async fn models(State(state): State<ServerState>) -> Json<Value> {
Json(model_list_payload(
state
.routes
.iter()
.map(|(model, entry)| (model.as_str(), entry.capabilities)),
))
}
The models() function extracts the route map from ServerState, iterating over each configured model and its associated ModelCapabilities. It passes this iterator to model_list_payload(), which constructs the final response body.
Building the Codex-Compatible Payload
The model_list_payload() function (lines 46–66 in lib.rs) assembles the dual-format response by mapping over the route entries twice:
fn model_list_payload<'a>(
entries: impl IntoIterator<Item = (&'a str, ModelCapabilities)>,
) -> Value {
let entries = entries.into_iter().collect::<Vec<_>>();
json!({
"object": "list",
"data": entries.iter()
.map(|(model, caps)| model_entry_json(model, *caps))
.collect::<Vec<_>>(),
"models": entries.iter()
.enumerate()
.map(|(priority, (model, caps))|
codex_model_entry_json(model, *caps, priority))
.collect::<Vec<_>>(),
"first_id": entries.first().map(|(m, _)| *m),
"last_id": entries.last().map(|(m, _)| *m),
"has_more": false,
"default_model": entries.first().map(|(m, _)| *m),
"model_pool": entries.iter().map(|(m, _)| *m).collect::<Vec<_>>(),
})
}
The function generates the data array using the generic model_entry_json() helper (lines 169–177), while the models array uses codex_model_entry_json() to inject Codex-specific fields. Both arrays share the same underlying route configuration but present different metadata schemas.
Codex-Specific Metadata Fields
The codex_model_entry_json() function (lines 1209–1248 in lib.rs) transforms ModelCapabilities into the Codex ModelInfo format:
fn codex_model_entry_json(model: &str,
capabilities: ModelCapabilities,
priority: usize) -> Value {
let tool_calling = capabilities.tool_calling.unwrap_or(true);
let reasoning = capabilities.reasoning.unwrap_or(false);
json!({
"slug": model,
"display_name": model,
"description": "Switchyard-routed model.",
"default_reasoning_level": if reasoning { json!("xhigh") } else { Value::Null },
"supported_reasoning_levels": if reasoning { reasoning_levels() } else { json!([]) },
"shell_type": if tool_calling { "shell_command" } else { "disabled" },
"visibility": "list",
"supported_in_api": true,
"priority": priority,
"additional_speed_tiers": [],
"base_instructions": "You are Codex, a coding agent.",
"supports_reasoning_summaries": reasoning,
"default_reasoning_summary": "none",
"support_verbosity": reasoning,
"default_verbosity": if reasoning { json!("low") } else { Value::Null },
"apply_patch_tool_type": if tool_calling { Some("freeform") } else { None },
"web_search_tool_type": "text",
"truncation_policy": {"mode": "tokens", "limit": 10_000},
"supports_parallel_tool_calls": tool_calling,
"supports_image_detail_original": false,
"context_window": capabilities.context_window,
"max_context_window": capabilities.context_window,
"effective_context_window_percent": 95,
"experimental_supported_tools": [],
"input_modalities": ["text"],
"supports_search_tool": false,
})
}
This function derives critical fields from the route's declared capabilities:
shell_type– Set to"shell_command"iftool_callingis enabled, otherwise"disabled".default_reasoning_level– Populated with"xhigh"only when thereasoningcapability is true.context_window– Mapped directly fromcapabilities.context_windowto bothcontext_windowandmax_context_window.
Route Discovery and Configuration Flow
Switchyard populates the /v1/models response through a configuration-driven pipeline:
- TOML Configuration – Each route defined in the Switchyard configuration file declares a model identifier and a
ModelCapabilitiesstruct (defined incrates/switchyard-server/src/config.rsat lines 96–108), specifying optional boolean flags fortool_callingandreasoning, plus numericcontext_windowvalues. - State Initialization – At startup, these routes register with
ServerState, making the model identifiers and capabilities available to the HTTP handlers. - Runtime Exposure – When a client queries
/v1/models, the endpoint handler consumesstate.routesto build the payload, ensuring the response reflects the current routing table without requiring a server restart.
Querying the Endpoint
You can inspect the Codex-compatible metadata using standard HTTP tools. The following curl command retrieves the full model list:
curl -s http://localhost:4000/v1/models | jq .
To access the Codex-specific fields programmatically, extract the models array from the response:
import httpx
resp = httpx.get("http://localhost:4000/v1/models")
payload = resp.json()
# Access the first Codex-compatible model entry
codex_entry = payload["models"][0]
print(f"Slug: {codex_entry['slug']}")
print(f"Context window: {codex_entry['context_window']}")
print(f"Shell type: {codex_entry['shell_type']}")
Summary
- Dual-format response – The
/v1/modelsendpoint returns both standard OpenAI-style model entries (data) and Codex-specific metadata (models) in a single payload. - Implementation location – The endpoint handler resides in
crates/switchyard-server/src/lib.rs(lines 60–66), with payload construction logic at lines 46–66. - Capability mapping –
codex_model_entry_json()transformsModelCapabilities(sourced from TOML configuration) into Codex-required fields likeshell_type,default_reasoning_level, andcontext_window. - Configuration-driven – Model capabilities defined in
config.rsflow throughServerStateto dynamically generate the response at runtime.
Frequently Asked Questions
What is the difference between the data and models fields in the Switchyard /v1/models response?
The data field contains a standard OpenAI-compatible model list using the model_entry_json() format, suitable for generic API clients. The models field contains Codex-specific metadata generated by codex_model_entry_json(), including fields like slug, shell_type, and base_instructions that Codex clients require to render ModelInfo cards and enable shell command features.
How does Switchyard determine the shell_type for a Codex model entry?
Switchyard sets shell_type to "shell_command" if the route's ModelCapabilities has tool_calling set to true (or defaults to true if unspecified). If tool calling is disabled, the value becomes "disabled". This logic executes inside codex_model_entry_json() in crates/switchyard-server/src/lib.rs.
Can I customize the Codex-specific metadata fields like base_instructions or truncation_policy?
Currently, codex_model_entry_json() uses hardcoded defaults for most Codex-specific fields, including base_instructions ("You are Codex, a coding agent.") and truncation_policy (10,000 token limit). To customize these values, you must modify the source code in crates/switchyard-server/src/lib.rs and rebuild the switchyard-server crate.
Which configuration file defines the capabilities used in the /v1/models response?
The capabilities exposed in the /v1/models response originate from the ModelCapabilities struct defined in crates/switchyard-server/src/config.rs (lines 96–108). These values are populated from the Switchyard TOML configuration file at startup and stored in ServerState, where the endpoint handler accesses them to build the JSON payload.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →