# How Agents Handle Failure Recovery and Escalation in the NEXUS Pipeline

> Discover how agents manage failure recovery and escalation in the NEXUS pipeline. Learn about the fail-fast, fix-fast system and automatic escalation for persistent issues.

- Repository: [Michael Sitarzewski/agency-agents](https://github.com/msitarzewski/agency-agents)
- Tags: internals
- Published: 2026-03-09

---

**The NEXUS pipeline implements a deterministic fail-fast, fix-fast system where tasks cycle through a Dev↔QA loop up to three times before the Agents Orchestrator automatically generates an Escalation Report and routes persistent failures to a designated gate-keeper such as the Studio Producer.**

The `msitarzewski/agency-agents` repository defines a strict failure management protocol within its NEXUS strategy. Every specialized agent operates under explicit retry limits and escalation procedures designed to prevent infinite loops while ensuring quality gates are met before production deployment.

## Core Failure Recovery Mechanics

### The Dev↔QA Validation Loop

Every task begins with a **Developer Agent** implementing the requirement, followed immediately by validation from the **Evidence Collector** (QA agent). This specialized QA agent runs visual and functional checks, returning a binary **PASS** or **FAIL** verdict. The loop is managed by the Agents Orchestrator, which tracks the state of each task in [`specialized/agents-orchestrator.md`](https://github.com/msitarzewski/agency-agents/blob/main/specialized/agents-orchestrator.md) under the *Quality Loop* section.

### Retry Limits with Feedback Integration

When QA returns **FAIL**, the orchestrator increments a retry counter stored in the task metadata. The system allows a **maximum of 3 attempts** per task. Each retry iteration includes the previous QA feedback—specific issues, screenshots, logs, and fix instructions—ensuring the developer agent has full context for the next attempt.

This logic is defined in [`specialized/agents-orchestrator.md`](https://github.com/msitarzewski/agency-agents/blob/main/specialized/agents-orchestrator.md) under *Failure Management*:

```yaml

# agents-orchestrator.md – Failure handling excerpt

failure_handling:
  max_retries: 3                     # ≤ 3 attempts per task

  on_fail:
    - increment_retry_counter
    - if_retry < max_retries:       # loop back to developer with feedback

        route_to: developer_agent
        include: qa_feedback
    - else:                         # 3rd failure → escalation

        generate: escalation_report
        route_to: studio_producer

```

### Automatic Escalation Triggers

After the third failed attempt (`retry >= 3`), the orchestrator automatically triggers the **Escalation Protocol**. At this stage, the task is no longer routed back to the original developer. Instead, the system generates a structured **Escalation Report** using the markdown template defined in [`strategy/nexus-strategy.md`](https://github.com/msitarzewski/agency-agents/blob/main/strategy/nexus-strategy.md) and forwards it to the **Studio Producer** or another designated gate-keeper.

## Escalation Protocol and Gate-Keeper Intervention

### Escalation Report Generation

The Escalation Report is a comprehensive markdown document that captures the full failure history, root-cause analysis, and recommended resolutions. According to [`strategy/nexus-strategy.md`](https://github.com/msitarzewski/agency-agents/blob/main/strategy/nexus-strategy.md), the report must include:

- **Failure History**: Chronological log of all three attempts with QA feedback
- **Root Cause Analysis**: Technical assessment of why the fix failed
- **Recommended Resolutions**: Specific actions such as re-assignment, task decomposition, or architecture revision
- **Impact Assessment**: Timeline and quality metrics affected by the blockage

### Gate-Keeper Decision Matrix

Once the Studio Producer receives the Escalation Report, they evaluate the issue using the decision matrix defined in the *Escalation Protocol* section of [`strategy/nexus-strategy.md`](https://github.com/msitarzewski/agency-agents/blob/main/strategy/nexus-strategy.md). The gate-keeper has four primary resolution paths:

1. **Re-assign**: Route the task to a different developer agent with specialized skills
2. **Decompose**: Split the complex task into smaller sub-tasks that can pass QA independently
3. **Revise Architecture**: Engage architect agents to adjust technical specifications or infrastructure
4. **Accept Limitation**: Mark the task as **BLOCKED** with documented constraints and move forward

The [`support/support-support-responder.md`](https://github.com/msitarzewski/agency-agents/blob/main/support/support-support-responder.md) file provides additional criteria for when to involve senior leadership or external resources, ensuring that business-critical blockages receive appropriate visibility.

## Configuration and Implementation Details

### QA Failure Feedback Template

When QA returns a failure, the feedback follows a strict template defined in [`strategy/nexus-strategy.md`](https://github.com/msitarzewski/agency-agents/blob/main/strategy/nexus-strategy.md) to ensure the developer receives actionable data:

```markdown

## QA Failure Feedback

### Task: [Task ID – e.g., FE-42]

### Attempt: 2/3

### Verdict: FAIL

#### Specific Issues Found

1. **Layout Shift**: Header overlaps content on mobile.
   - Expected: Fixed header stays above content.
   - Actual: Header pushes content down.
   - Evidence: `header_shift.png`

2. **API Timeout**: `/api/v1/items` returns 504 after 30 s.
   - Expected: ≤ 200 ms response.
   - Actual: 30 s timeout.
   - Evidence: `api_timeout.log`

#### Fix Instructions

- Adjust CSS `position: sticky` on header.
- Optimize DB query for `/api/v1/items` (add index on `status`).

#### Files to Modify

- `src/components/Header.vue`
- `src/api/items.js`

#### Retry Expectations

- Implement fixes only; do **not** add new features.
- Submit for QA as attempt **3/3`.

```

### Escalation Report Template

The final escalation document used by the Studio Producer follows this structure from [`strategy/nexus-strategy.md`](https://github.com/msitarzewski/agency-agents/blob/main/strategy/nexus-strategy.md):

```markdown

## Escalation Report

### Task: FE-42 – Mobile Header Layout

### Attempts Exhausted: 3/3

### Escalation Level: To Studio Producer

#### Failure History

- **Attempt 1**: Header overlaps; no visual proof.
- **Attempt 2**: Same issue; added screenshots.
- **Attempt 3**: Still fails after CSS refactor.

#### Root Cause Analysis

- CSS `position: sticky` not compatible with current layout grid.
- Backend API latency due to missing index.

#### Recommended Resolution

- [ ] Reassign to a Senior Frontend Developer.
- [ ] Decompose UI fix into two sub‑tasks (header, grid).
- [ ] backend team adds index on `status`.
- [ ] Schedule a sprint‑level design review.

#### Impact Assessment

- **Blocking**: Subsequent UI screens depend on fixed header.
- **Timeline Impact**: +2 days to sprint.
- **Quality Impact**: Current QA pass rate drops to 68 %.

```

## Summary

- The **Agents Orchestrator** enforces a **3-attempt limit** per task, cycling failed work through a Dev↔QA feedback loop defined in [`specialized/agents-orchestrator.md`](https://github.com/msitarzewski/agency-agents/blob/main/specialized/agents-orchestrator.md).
- **QA failures** trigger detailed feedback reports that include specific issues, evidence files, and fix instructions to guide the next retry attempt.
- After the third failure, the system automatically generates an **Escalation Report** and routes the task to a **gate-keeper** such as the Studio Producer.
- The **gate-keeper** can re-assign, decompose, revise architecture, or accept limitations based on the escalation protocol in [`strategy/nexus-strategy.md`](https://github.com/msitarzewski/agency-agents/blob/main/strategy/nexus-strategy.md).
- **Support Responders** provide additional routing criteria for involving senior leadership or external resources when business-critical blockages occur.

## Frequently Asked Questions

### What happens when a task fails QA three times?

After the third QA failure, the Agents Orchestrator stops the retry loop and triggers the Escalation Protocol. The system generates a structured Escalation Report containing the full failure history, root-cause analysis, and recommended resolutions, then routes the task to the Studio Producer or another designated gate-keeper for manual intervention.

### Who receives the Escalation Report and what can they do?

The **Studio Producer** typically receives the Escalation Report, though the routing may vary based on criteria in [`support/support-support-responder.md`](https://github.com/msitarzewski/agency-agents/blob/main/support/support-support-responder.md). The gate-keeper has four primary resolution options: re-assign the task to a different developer agent, decompose the task into smaller sub-tasks, revise the technical architecture through architect agents, or accept the limitation and mark the task as blocked.

### How does the QA feedback loop improve failure recovery?

Each retry attempt includes a detailed **QA Failure Feedback** template that specifies exact issues found, evidence files (screenshots, logs), expected vs. actual behavior, and precise fix instructions. This ensures the developer agent receives actionable context rather than generic error messages, maximizing the probability of success on subsequent attempts within the three-retry limit.

### Can retry limits be configured for different task types?

The default configuration in [`specialized/agents-orchestrator.md`](https://github.com/msitarzewski/agency-agents/blob/main/specialized/agents-orchestrator.md) sets a universal maximum of **3 retries** per task before mandatory escalation. While the source documentation emphasizes this consistent limit to maintain pipeline velocity, specific task types or critical paths could theoretically implement custom retry logic by modifying the `failure_handling` configuration block in the orchestrator's definition file.