How no-mistakes Implements Agent Fallbacks for Reliable AI Operations

The no-mistakes framework handles agent fallbacks through a transparent retry mechanism in internal/agent/fallback.go that automatically switches to alternate agents when the primary AI provider is unavailable, ensuring deterministic pipeline execution even when language-model servers fail.

The no-mistakes repository provides an AI-powered automation framework where critical steps like code review, documentation generation, and linting are executed through specialized agents. Because these agents depend on external language-model servers that may crash or fail to start, the framework implements robust agent fallbacks that wrap multiple providers and guarantee pipeline completion without manual intervention.

Core Fallback Architecture in internal/agent/fallback.go

The fallback system centers on a fallbackAgent struct that implements the same Agent interface as concrete providers. This wrapper maintains a slice of agent implementations and delegates operations according to availability and capability requirements.

Construction and Initialization

The NewFallback constructor (lines 14-26) creates the appropriate wrapper based on the input slice size. When passed an empty slice, it returns a no-op implementation. A single-element slice returns the agent directly to eliminate unnecessary overhead. Only multiple agents trigger the creation of a fallbackAgent that owns a defensive copy of the slice, ensuring the underlying array cannot be modified by callers.

Session-Aware Routing

When executing operations, the wrapper examines RunOpts to determine session requirements. Lines 74-84 implement a strict filtering mechanism: if RunOpts contains a non-empty Session, the wrapper restricts execution to the single agent that can provide that specific session. This prevents session mixing across different providers, maintaining state consistency. Otherwise, the wrapper iterates through all configured agents in the order specified during construction.

Capability Aggregation

The wrapper implements capability queries by aggregating results across all wrapped agents. As implemented in lines 36-71, methods like SupportsSessionResume return the logical OR of all agents' capabilities, while NeutralizesGateInstructions uses logical AND to ensure safety. This design guarantees that the fallback wrapper only reports a capability when it is uniformly safe or explicitly supported by at least one member.

The Fallback Execution Loop

The Run method (lines 90-113) implements the core retry logic that distinguishes transient failures from permanent errors. For each candidate agent, the wrapper performs four distinct operations:

  1. Option Adjustment: If the current candidate cannot resume a session, the session fields are cleared from RunOpts (lines 90-93).
  2. Execution: The wrapper invokes current.Run and emits an attempt record when the agent does not self-report attempts (lines 95-98).
  3. Success Handling: On successful completion, the result propagates immediately and the Provider field is populated if missing (lines 99-103).
  4. Error Classification: On failure, the wrapper checks if the error represents an unavailable agent using isAgentUnavailableError (lines 30-47). If unavailable and additional candidates exist, the wrapper sends a fallback message via opts.OnChunk and continues iteration (lines 105-113). Otherwise, it returns the error immediately.

Error Classification and Retry Logic

The isAgentUnavailableError helper (lines 30-47) identifies catastrophic failures by detecting substrings that indicate the agent process never started or exited unexpectedly. This distinguishes infrastructure failures from generation errors that should not trigger fallback. The companion fallbackReason function (lines 50-61) truncates verbose error messages for user-facing output, ensuring diagnostic clarity without exposing internal stack traces.

User-Facing Messaging

During fallback transitions, the wrapper communicates state changes through the OnChunk callback defined in RunOpts. This mechanism prints concise messages like "agent X failed ... falling back to Y", giving users visibility into the retry sequence without interrupting pipeline execution.

Pipeline Integration and Static Fallbacks

The fallback mechanism integrates throughout the no-mistakes pipeline. In internal/pipeline/steps/pr.go, the framework attempts to generate PR content through an agent, but if that call fails, it falls back to a deterministic PR body via fallbackPRContent (lines 179-180). Similar patterns appear in internal/pipeline/steps/document.go for documentation generation and internal/pipeline/steps/common_fix.go for commit message summaries. These implementations demonstrate how agent fallbacks provide graceful degradation when AI generation is impossible.

Implementation Examples

Creating a fallback agent requires instantiating concrete providers and wrapping them in order of preference:

// Choose the primary and secondary agents (both implement the Agent interface)
primary := codex.NewAgent(cfg)
secondary := claude.NewAgent(cfg)

// The fallback wrapper will try primary first, then secondary on unavailability.
agent := agent.NewFallback([]agent.Agent{primary, secondary})

Running a step with fallback transparency:

opts := agent.RunOpts{
    Prompt:  "Write a concise PR description for the change.",
    OnChunk: func(msg string) { fmt.Print(msg) }, // prints fallback messages
}
result, err := agent.Run(context.Background(), opts)
if err != nil {
    // All agents failed – handle the error or use a static fallback.
    log.Fatalf("no agent could produce a result: %v", err)
}
fmt.Println("Agent result:", result.Text)

Typical usage inside pipeline steps follows this pattern from internal/pipeline/steps/pr.go:

result, err := agents.Run(ctx, agent.RunOpts{
    Prompt:  prPrompt,
    OnChunk: func(msg string) { log.Info(msg) },
})
if err != nil {
    // Log the failure and use a deterministic PR body.
    slog.Warn("agent failed for PR content, using fallback", "error", err)
    return fallbackPRContent(...)
}
return result.Text, nil

Summary

  • Deterministic ordering: Agents execute in the sequence supplied to NewFallback, ensuring predictable behavior across runs.
  • Session isolation: When RunOpts specifies a session, only the owning agent executes, preventing cross-provider state contamination.
  • Capability aggregation: The wrapper reports capabilities using logical OR for availability features and logical AND for safety-critical properties.
  • Graceful degradation: Infrastructure failures trigger automatic retry with alternate agents, while generation errors propagate immediately for proper handling.
  • Resource cleanup: The Close method (lines 17-27) iterates through all wrapped agents and aggregates close errors into a single combined error.

Frequently Asked Questions

How does no-mistakes determine when to trigger an agent fallback?

The framework triggers fallbacks only for infrastructure-level failures detected by isAgentUnavailableError in internal/agent/fallback.go. This function examines error messages for substrings indicating the agent process never started or crashed unexpectedly. Generation errors, content policy violations, or invalid outputs do not trigger fallbacks and propagate immediately to the caller.

Can I mix agents with different session capabilities in a single fallback chain?

Yes, but with constraints. The fallbackAgent wrapper supports heterogeneous agents through capability aggregation. However, if a RunOpts specifies a session, the wrapper restricts execution to the single agent that can provide that session (lines 74-84). This ensures session continuity while still allowing the fallback chain to include agents with varying capabilities for stateless operations.

What happens if all agents in the fallback chain fail?

If every candidate agent fails with an unavailable error, the wrapper returns the last encountered error unchanged (lines 105-113). This preserves the original failure context for diagnostic purposes. Pipeline steps like those in internal/pipeline/steps/pr.go then catch this error and execute deterministic fallback logic, such as generating static PR content, ensuring the pipeline completes even when AI agents are entirely inaccessible.

How does the fallback wrapper handle resource cleanup?

The Close method iterates over all wrapped agents and aggregates any close errors into a single combined error (lines 17-27). This ensures that shutting down a fallbackAgent properly releases resources across all underlying providers, regardless of which agents were actually invoked during execution.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →