How Q00/ouroboros PAL Router Tier System Handles Auto-Escalation and Auto-Downgrade
The Q00/ouroboros PAL Router implements auto-escalation by promoting tasks to higher tiers after two consecutive failures tracked by the EscalationManager, while auto-downgrade occurs implicitly through stateless complexity-score recalculation that resets when tasks succeed.
The Q00/ouroboros repository implements a Progressive Adaptive LLM (PAL) router that dynamically assigns tasks to three cost-based tiers—Frugal, Standard, and Frontier—based on runtime complexity and failure history. This system optimizes API costs by starting with cheaper models and automatically escalating only when tasks prove too difficult for lower tiers. Understanding the interaction between the static complexity thresholds and the dynamic failure tracker is essential for configuring resilient model orchestration pipelines.
Stateless Auto-Downgrade via Complexity Score Routing
The router's auto-downgrade capability stems from its stateless design, where every task is evaluated fresh against complexity thresholds rather than maintaining persistent tier assignments.
Complexity Estimation and Thresholds
In src/ouroboros/routing/complexity.py, the estimate_complexity function generates a ComplexityScore between 0.0 and 1.0. The src/ouroboros/routing/router.py file defines static thresholds at lines 52-54:
# src/ouroboros/routing/router.py
THRESHOLD_FRUGAL = 0.4
THRESHOLD_STANDARD = 0.7
The _select_tier_from_score method (lines 78-91) maps these scores to tiers:
def _select_tier_from_score(score: float) -> Tier:
if score < THRESHOLD_FRUGAL:
return Tier.FRUGAL
if score < THRESHOLD_STANDARD:
return Tier.STANDARD
return Tier.FRONTIER
Because the router recalculates the tier for every task without maintaining historical state, a previously escalated pattern automatically reverts to Frugal or Standard when subsequent tasks present lower complexity scores. This implicit auto-downgrade requires no explicit downgrade logic—the system simply routes based on the current task's characteristics.
Failure-Based Auto-Escalation via EscalationManager
While the router itself is stateless, the EscalationManager in src/ouroboros/routing/escalation.py maintains failure counters to drive vertical tier promotion when tasks repeatedly fail.
Failure Tracking and Thresholds
The manager tracks consecutive failures per pattern using the FailureTracker class (lines 48-60). Auto-escalation triggers after FAILURE_THRESHOLD = 2 consecutive failures (line 44).
Escalation Flow and Tier Promotion
When record_failure is called, the manager increments the counter and returns an EscalationAction indicating whether to promote the tier:
from ouroboros.routing.escalation import EscalationManager
from ouroboros.routing.tiers import Tier
manager = EscalationManager()
pattern = "deploy_app"
# First failure – counter increments, no escalation yet
action = manager.record_failure(pattern, Tier.FRUGAL).value
assert not action.should_escalate
# Second consecutive failure – escalation to Standard triggered
action = manager.record_failure(pattern, Tier.FRUGAL).value
assert action.should_escalate
assert action.target_tier == Tier.STANDARD
The _get_next_tier method (lines 62-75) implements the promotion path: Frugal → Standard → Frontier.
Stagnation Detection at Frontier
If failures persist at the Frontier tier, the system emits a stagnation event rather than attempting further vertical escalation. This prevents infinite retries on inherently unsolvable tasks and signals the resilience layer to attempt lateral-thinking strategies.
Success-Based Counter Reset
The record_success method (lines 74-84) clears the failure counter while preserving the current tier for the active task. This reset allows subsequent routing decisions to potentially downgrade based on fresh complexity scoring:
# Success resets counter – next failure starts from fresh state
manager.record_success(pattern)
action = manager.record_failure(pattern, Tier.STANDARD).value
assert not action.should_escalate # counter reset, need two failures again
Runtime Integration Pattern
The orchestration layer in src/ouroboros/plugin/orchestration/router.py integrates these components by coordinating the stateless router with the stateful escalation tracker:
- The orchestrator calls the PAL router for initial tier selection based on
estimate_complexity. - The task executes using the selected tier (Frugal, Standard, or Frontier).
- On failure, the orchestrator calls
EscalationManager.record_failure; on success, it callsrecord_success. - If escalation is triggered, the orchestrator re-routes the task with the upgraded tier returned in the
EscalationAction.
Practical Implementation Examples
Basic Routing with Auto-Downgrade Potential
from ouroboros.routing.router import route_task
from ouroboros.routing.complexity import TaskContext
from ouroboros.routing.tiers import Tier
# Low-complexity task routes to Frugal tier
ctx = TaskContext(token_count=120, tool_dependencies=[], ac_depth=1)
result = route_task(ctx)
assert result.is_ok and result.value.tier == Tier.FRUGAL
# Subsequent high-complexity task escalates based on score
ctx_high = TaskContext(token_count=5000, tool_dependencies=["database"], ac_depth=5)
result_high = route_task(ctx_high)
assert result_high.value.tier == Tier.FRONTIER
Failure-Driven Escalation Sequence
from ouroboros.routing.escalation import EscalationManager
from ouroboros.routing.tiers import Tier
manager = EscalationManager()
pattern = "complex_analysis"
# Simulate two failures at Frugal tier
manager.record_failure(pattern, Tier.FRUGAL)
action = manager.record_failure(pattern, Tier.FRUGAL).value
# Verify escalation to Standard
assert action.should_escalate
assert action.target_tier == Tier.STANDARD
Handling Frontier Stagnation
# After escalating to Frontier and encountering repeated failures
manager.record_failure("heavy_query", Tier.FRONTIER)
action = manager.record_failure("heavy_query", Tier.FRONTIER).value
# Stagnation event emitted instead of further escalation
assert action.is_stagnation
assert action.target_tier is None
# Create event for resilience system
event = manager.create_stagnation_event("heavy_query", action.failure_count)
Summary
- Auto-escalation is driven by the
EscalationManagerinsrc/ouroboros/routing/escalation.py: after two consecutive failures (FAILURE_THRESHOLD = 2), a pattern is automatically promoted to the next tier following the Frugal → Standard → Frontier path. - Auto-downgrade happens implicitly because the PAL router in
src/ouroboros/routing/router.pyis stateless and re-computes the tier from the current task's complexity score on every invocation. - When the Frontier tier still fails after the threshold, the system emits a stagnation event to trigger lateral resilience strategies rather than infinite vertical escalation.
- The success reset mechanism (
record_success) clears failure counters immediately, allowing the system to fall back to cheaper tiers for subsequent tasks.
Frequently Asked Questions
What triggers auto-escalation in the Q00/ouroboros PAL Router?
Auto-escalation triggers when the EscalationManager detects two consecutive failures for the same pattern at a given tier. The FAILURE_THRESHOLD constant (set to 2 in src/ouroboros/routing/escalation.py line 44) determines this boundary, after which the system returns an EscalationAction with should_escalate=True and the next tier target.
How does the system automatically downgrade to cheaper tiers?
The system implements auto-downgrade implicitly through stateless routing. Since _select_tier_from_score in src/ouroboros/routing/router.py recalculates the tier based on the current task's complexity score (0.0-1.0) without considering historical tier assignments, any task with a score below 0.4 routes to Frugal and below 0.7 routes to Standard, regardless of previous escalations.
What happens when the Frontier tier fails repeatedly?
When the Frontier tier fails twice consecutively, the EscalationManager emits a stagnation event instead of attempting further escalation. This signals to the orchestration layer that vertical tier promotion has reached its limit and alternative strategies (such as retry with different prompts or human intervention) should be attempted.
Where are the tier thresholds defined in the codebase?
The tier thresholds are defined as constants in src/ouroboros/routing/router.py at lines 52-54: THRESHOLD_FRUGAL = 0.4 and THRESHOLD_STANDARD = 0.7. These values are used by the _select_tier_from_score method (lines 78-91) to map complexity scores to the Frugal, Standard, or Frontier tiers.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →