# How Agent Fallback Works in no-mistakes: Resilient AI Agent Orchestration

> Discover how no-mistakes agent fallback ensures resilient AI orchestration. Learn how the fallbackAgent wrapper transparently retries operations preserving session integrity and surfacing concise messages on chunk.

- Repository: [Kun Chen/no-mistakes](https://github.com/kunchenguid/no-mistakes)
- Tags: deep-dive
- Published: 2026-07-19

---

**In no-mistakes, agent fallback is implemented by the `fallbackAgent` wrapper in [`internal/agent/fallback.go`](https://github.com/kunchenguid/no-mistakes/blob/main/internal/agent/fallback.go), which transparently retries operations across a prioritized list of concrete agents when the primary agent is unavailable, while preserving session integrity and surfacing concise retry messages via `OnChunk`.**

The no-mistakes project runs AI-powered steps like code review, documentation generation, and linting through pluggable agents. When an underlying language-model server fails to start or crashes mid-operation, the framework needs a deterministic way to recover. Its agent fallback mechanism handles this transparently, ensuring that pipeline steps degrade gracefully rather than fail outright.

## Fallback Agent Construction and Interface

At the heart of the system is **[`internal/agent/fallback.go`](https://github.com/kunchenguid/no-mistakes/blob/main/internal/agent/fallback.go)**. The `NewFallback` constructor accepts a slice of `Agent` implementations and returns an appropriate strategy based on the input size: a no-op agent for an empty slice, the single agent directly when only one is provided, or a `fallbackAgent` that owns a copy of the slice (lines 14–26). This wrapper exposes the same `Agent` interface as its delegates, making it transparent to callers.

The wrapper’s `Name` method simply delegates to the first configured agent (lines 29–34), ensuring that telemetry and logging identify the primary candidate by default.

## Capability Aggregation and Session Handling

### Querying Agent Capabilities

Because the fallback wrapper must speak for multiple backends, it aggregates capability checks across every wrapped agent. Methods such as `SupportsSessionResume`, `SupportsSessionProvider`, and `NeutralizesGateInstructions` iterate the inner agents and return the logical OR or AND of their individual results (lines 36–71). The wrapper reports `SupportsSessionResume` as true if **any** inner agent supports it, while it only reports `NeutralizesGateInstructions` when **all** members satisfy that condition. This prevents the fallback group from advertising unsafe or partially implemented behaviors.

### Session-Aware Routing

When `RunOpts` contains a non-empty `Session`, the wrapper restricts execution to the single agent that can provide that exact session (lines 74–84). If no such agent exists, the session is not forced onto an incompatible backend. This design guarantees that session state is never mixed across different providers.

## The Run Loop and Error Classification

### Attempt, Retry, and User Messaging

The `Run` method implements a sequential retry loop over the candidate agents. For each iteration it performs four main operations:

1. **Sanitize options:** If the current candidate cannot resume a session, the wrapper clears the session fields in `RunOpts` before invoking it (lines 90–93).
2. **Execute and record:** It calls `current.Run` and, when the agent does not self-report attempts, emits an attempt record automatically (lines 95–98).
3. **Propagate success:** On a successful result, the wrapper returns the output immediately and backfills the `Provider` field if the agent left it empty (lines 99–103).
4. **Handle failure:** On error, the wrapper stores the error and checks whether it signals agent unavailability. If `isAgentUnavailableError` returns true and another candidate remains, the wrapper sends a concise fallback message through `opts.OnChunk` and continues the loop (lines 105–113). If the error is not an availability issue or no candidates remain, the error is returned immediately.

### Detecting Unavailability and Formatting Reasons

The helper `isAgentUnavailableError` inspects error strings for substrings that indicate the agent process never started or exited unexpectedly (lines 30–47). Meanwhile, `fallbackReason` truncates lengthy internal error messages into short, user-facing strings (lines 50–61), keeping console output readable without hiding the fact that a failover occurred.

## Graceful Shutdown and Pipeline Integration

### Closing All Wrapped Agents

The `Close` method iterates over every wrapped agent and aggregates any shutdown errors into a single combined error (lines 17–27). This ensures that resources for all backends are released even when earlier agents in the chain have already failed.

### Fallback in Pipeline Steps

The fallback mechanism is exercised throughout the no-mistakes pipeline. In [`internal/pipeline/steps/pr.go`](https://github.com/kunchenguid/no-mistakes/blob/main/internal/pipeline/steps/pr.go), for instance, the step invokes an agent to generate PR content; if that call fails, it falls back to a deterministic PR body through `fallbackPRContent` (lines 179–180). Similar patterns appear in [`internal/pipeline/steps/document.go`](https://github.com/kunchenguid/no-mistakes/blob/main/internal/pipeline/steps/document.go), lint steps, and test steps, all of which guarantee a deterministic outcome when an AI agent cannot produce a result.

```go
// Simplified excerpt from internal/pipeline/steps/pr.go
result, err := agents.Run(ctx, agent.RunOpts{
    Prompt:  prPrompt,
    OnChunk: func(msg string) { log.Info(msg) },
})
if err != nil {
    slog.Warn("agent failed for PR content, using fallback", "error", err)
    return fallbackPRContent(...)
}
return result.Text, nil

```

## Creating a Fallback Agent in Practice

To leverage this mechanism, instantiate your primary and secondary agents and pass them to `NewFallback` in priority order:

```go
primary := codex.NewAgent(cfg)
secondary := claude.NewAgent(cfg)

agent := agent.NewFallback([]agent.Agent{primary, secondary})

```

Running the composite agent is identical to running any other agent:

```go
opts := agent.RunOpts{
    Prompt:  "Write a concise PR description for the change.",
    OnChunk: func(msg string) { fmt.Print(msg) },
}
result, err := agent.Run(context.Background(), opts)
if err != nil {
    log.Fatalf("no agent could produce a result: %v", err)
}
fmt.Println("Agent result:", result.Text)

```

Because `fallbackAgent` conforms to the standard `Agent` interface, consumers in `internal/pipeline/steps` require no special logic to benefit from resilient retries.

## Summary

- **[`internal/agent/fallback.go`](https://github.com/kunchenguid/no-mistakes/blob/main/internal/agent/fallback.go)** contains the core `fallbackAgent` wrapper that implements agent fallback in no-mistakes.
- `NewFallback` constructs a transparent wrapper that preserves the `Agent` interface for all callers.
- Capability methods aggregate results across wrapped agents using logical OR/AND to ensure safe feature advertisement.
- Session affinity is enforced by routing session-bearing requests only to the agent that owns the session.
- Unavailability errors trigger sequential retries with concise user-visible messages via `opts.OnChunk`.
- Pipeline steps such as PR generation in [`internal/pipeline/steps/pr.go`](https://github.com/kunchenguid/no-mistakes/blob/main/internal/pipeline/steps/pr.go) rely on this wrapper to degrade gracefully to static fallback content.

## Frequently Asked Questions

### How does no-mistakes decide when to trigger a fallback agent?

The `fallbackAgent` loop classifies each error using `isAgentUnavailableError`, which looks for substrings indicating that the agent process never started or exited unexpectedly. Only availability-related errors trigger a retry with the next candidate; all other errors are returned immediately to preserve their semantics.

### Can I mix agents that support sessions with agents that do not?

Yes, but the wrapper prevents unsafe mixing. If `RunOpts` includes a non-empty `Session`, the wrapper restricts the candidate list to the single agent that can resume that specific session. For candidates that do not support session resume, the wrapper temporarily clears session fields before invoking them.

### What happens if every agent in the fallback list fails?

If all candidates fail, the last encountered error is returned unchanged. This preserves the original failure cause for diagnostics while ensuring that callers upstream—such as pipeline steps in [`internal/pipeline/steps/pr.go`](https://github.com/kunchenguid/no-mistakes/blob/main/internal/pipeline/steps/pr.go)—can implement their own static fallback logic if desired.

### Does the fallback wrapper affect how capabilities are reported to the rest of the framework?

Yes. The wrapper aggregates capability queries across all inner agents. Capabilities like `SupportsSessionResume` return true if **any** wrapped agent supports them, whereas `NeutralizesGateInstructions` returns true only if **all** wrapped agents satisfy it. This keeps the framework’s capability checks accurate when multiple backends are involved.