How to Optimize LLM Request Timeout and Performance Tuning in AxonHub
To optimize LLM request timeout and performance tuning in AxonHub, configure the hierarchical timeout layers in conf/conf.go—set server.request_timeout for HTTP bounds, server.llm_request_timeout for orchestration limits, and provider-specific timeouts to match your workload latency requirements.
AxonHub manages LLM requests through a sophisticated hierarchy of timeouts and performance controls that govern everything from HTTP server bounds to provider-level network calls. Understanding how to optimize LLM request timeout and performance tuning across these layers ensures low-latency responses while preventing cascading failures under heavy load.
Understanding AxonHub's Timeout Architecture
AxonHub implements a five-layer timeout hierarchy that governs different stages of request processing. Each layer serves a distinct purpose and can be tuned independently based on your deployment characteristics.
The HTTP server timeout acts as the outermost boundary, controlling the maximum duration for the entire inbound HTTP request including authentication, routing, and response marshaling. The LLM request timeout governs the orchestration pipeline itself, covering model selection, load-balancing, and streaming operations. The HTTP client timeout manages network-level calls to upstream providers like OpenAI or Anthropic. Finally, token exchange and cache refresh timeouts handle background authentication and metadata operations.
Configuring Server-Level Timeouts
HTTP Server Request Timeout
The HTTP server timeout is defined in conf/conf.go within the ServerConfig struct as the RequestTimeout field. This value represents the maximum time the Gin server will wait for the entire request lifecycle.
By default, this is set to 30s via the server.request_timeout configuration key. You can override this using the environment variable AXONHUB_SERVER_REQUEST_TIMEOUT (accepting values like 30s, 1m, or 2m) or by modifying the YAML/JSON configuration file directly. The timeout is applied to the Gin router in internal/server/server.go.
LLM Orchestration Timeout
The LLM request timeout controls how long the orchestration pipeline may spend processing a single request, including model selection, load balancing, and streaming generation. This is defined in conf/conf.go as LLMRequestTimeout on the ServerConfig struct.
The default value is 600s (10 minutes), configurable via the server.llm_request_timeout key or the AXONHUB_SERVER_LLM_REQUEST_TIMEOUT environment variable. This timeout is propagated through the request context in internal/server/orchestrator/request_execution.go, ensuring that all downstream operations respect the deadline.
Tuning Provider and Internal Timeouts
Provider HTTP Client Timeouts
Each LLM provider implementation maintains its own HTTP client timeout for network calls to upstream services like OpenAI, Anthropic, or Vertex AI. These are typically hard-coded to 30s in provider files such as axon/provider/openai/provider.go.
You can override these via provider-specific configuration keys such as provider.openai.timeout. Ensure that provider timeouts are set greater than or equal to the LLM request timeout to prevent the client from aborting before the orchestration completes.
Token Exchange Timeouts
Authentication operations, particularly OAuth and API-key token refreshes, are governed by timeouts defined in llm/oauth/exchange_strategy.go. The default implementation uses context.WithTimeout(parent, 30*time.Second).
Increase this value only when using slow authentication flows or when experiencing transient network glitches during token refresh. You can modify the constant directly in the source file or expose a configuration flag if your deployment requires longer refresh windows.
Cache Refresh Timeouts
Background cache operations for model metadata and pricing information use a refresh timeout defined in internal/pkg/xcache/live/cache.go. The default RefreshTimeout is 30s.
For deployments with thousands of models, increase this via the cache.refresh_timeout configuration key to prevent the background loader from being cancelled before completing metadata retrieval.
Performance Tuning Best Practices
Follow these six guidelines to optimize LLM request timeout and performance tuning for your specific workload:
-
Start with the LLM request timeout – Set this just long enough for your longest expected generation (e.g.,
2mfor 8k token outputs). Shorter values reduce latency but surface "request timed out" errors for larger outputs. -
Match the HTTP client timeout – Ensure provider network timeouts are greater than or equal to the LLM request timeout to prevent premature client abortion.
-
Align server request timeout – Keep the HTTP server timeout less than or equal to the LLM request timeout (e.g.,
server.request_timeout = 3m,server.llm_request_timeout = 5m) to avoid 504 Gateway Timeout errors. -
Increase token exchange timeout only for slow auth – Most providers refresh tokens in under 5 seconds; raise this value only when using device flows or experiencing network instability.
-
Leverage cache refresh timeout for large catalogs – If hosting thousands of models, increase
cache.refresh_timeoutto2mor higher to allow background metadata loading to complete. -
Profile the pipeline – Use the built-in tracing in
internal/tracing/log.goto identify latency bottlenecks in model selection, load-balancing, or provider streaming, then adjust load-balancer strategies accordingly.
Programmatic Configuration Example
You can override timeout values programmatically at startup by modifying the configuration struct before initializing the server:
package main
import (
"time"
"github.com/looplj/axonhub/conf"
"github.com/looplj/axonhub/internal/server"
)
func main() {
// Load default config (reads from /conf.yaml, env vars, etc.)
cfg := conf.Load()
// Override timeouts programmatically
cfg.Server.RequestTimeout = 3 * time.Minute // overall HTTP request
cfg.Server.LLMRequestTimeout = 5 * time.Minute // orchestration timeout
cfg.Provider.OpenAI.HTTPTimeout = 4 * time.Minute // provider HTTP client
cfg.Cache.RefreshTimeout = 2 * time.Minute // background cache refresh
// Initialise the server with the customised config
srv := server.New(cfg)
srv.Run()
}
Source references: conf/conf.go defines the default values; internal/server/server.go applies the HTTP timeout to the Gin router; provider implementations such as axon/provider/openai/provider.go configure the HTTP client timeout.
Key Source Files for Timeout Management
conf/conf.go– DefinesServerConfig.RequestTimeoutandServerConfig.LLMRequestTimeoutwith default values and configuration mappings.internal/server/server.go– Applies the HTTP server timeout to the Gin router and initializes the server with these bounds.internal/server/orchestrator/request_execution.go– Propagates the LLM request timeout throughcontext.Contextto all downstream operations.axon/provider/openai/provider.go(and similar provider files) – Configures provider-specific HTTP client timeouts for upstream API calls.llm/oauth/exchange_strategy.go– Manages the 30-second timeout for OAuth and API-key token refresh operations.internal/pkg/xcache/live/cache.go– Controls theRefreshTimeoutfor background model metadata and pricing cache updates.internal/tracing/log.go– Provides latency tracing to identify performance bottlenecks across the request pipeline.
Summary
- AxonHub uses a five-layer timeout hierarchy: HTTP server, LLM orchestration, provider HTTP client, token exchange, and cache refresh.
- Configure
server.request_timeoutandserver.llm_request_timeoutinconf/conf.goor via environment variablesAXONHUB_SERVER_REQUEST_TIMEOUTandAXONHUB_SERVER_LLM_REQUEST_TIMEOUT. - Ensure provider HTTP timeouts exceed LLM request timeouts to prevent premature connection drops during long generations.
- Adjust cache refresh timeouts when hosting large model catalogs to prevent background loader cancellation.
- Use programmatic configuration at startup to dynamically set timeouts based on deployment environment.
- Leverage tracing in
internal/tracing/log.goto profile latency across model selection, load-balancing, and provider calls.
Frequently Asked Questions
What is the difference between server request timeout and LLM request timeout in AxonHub?
The server request timeout (server.request_timeout) controls the maximum duration for the entire HTTP request lifecycle handled by the Gin server, including authentication, routing, and response marshaling. The LLM request timeout (server.llm_request_timeout) governs specifically the orchestration pipeline duration, covering model selection, load-balancing, and streaming generation. According to the source code in conf/conf.go, the server timeout defaults to 30 seconds while the LLM timeout defaults to 600 seconds.
How do I prevent 504 Gateway Timeout errors when processing large LLM outputs?
To prevent 504 errors, ensure the HTTP server request timeout is less than or equal to the LLM request timeout, and verify that both values exceed your longest expected generation time. For example, set server.request_timeout = 3m and server.llm_request_timeout = 5m in your configuration. Additionally, ensure provider HTTP client timeouts (configured in files like axon/provider/openai/provider.go) are set greater than or equal to the LLM request timeout to prevent the upstream connection from closing prematurely.
Where should I configure timeouts for specific LLM providers like OpenAI?
Provider-specific timeouts are configured within each provider's implementation file, such as axon/provider/openai/provider.go for OpenAI or axon/provider/anthropic/provider.go for Anthropic. These files instantiate http.Client with hard-coded default timeouts of 30 seconds. You can override these via provider-specific configuration keys (e.g., provider.openai.timeout) to ensure the network timeout accommodates long-running generations without interrupting the orchestration pipeline.
How does AxonHub handle context propagation across timeout layers?
AxonHub uses Go's context.Context to propagate timeout deadlines through the entire request tree. The LLMRequestTimeout defined in conf/conf.go is injected into the context in internal/server/orchestrator/request_execution.go, which then propagates to all downstream operations including provider calls and cache lookups. The internal/pkg/xcontext/context.go file provides helper functions like DetachWithTimeout to ensure spawned goroutines respect the parent request deadline while maintaining timeout isolation.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →