How Agents Handle Failure Recovery and Escalation in the NEXUS Pipeline
The NEXUS pipeline implements a deterministic fail-fast, fix-fast system where tasks cycle through a Dev↔QA loop up to three times before the Agents Orchestrator automatically generates an Escalation Report and routes persistent failures to a designated gate-keeper such as the Studio Producer.
The msitarzewski/agency-agents repository defines a strict failure management protocol within its NEXUS strategy. Every specialized agent operates under explicit retry limits and escalation procedures designed to prevent infinite loops while ensuring quality gates are met before production deployment.
Core Failure Recovery Mechanics
The Dev↔QA Validation Loop
Every task begins with a Developer Agent implementing the requirement, followed immediately by validation from the Evidence Collector (QA agent). This specialized QA agent runs visual and functional checks, returning a binary PASS or FAIL verdict. The loop is managed by the Agents Orchestrator, which tracks the state of each task in specialized/agents-orchestrator.md under the Quality Loop section.
Retry Limits with Feedback Integration
When QA returns FAIL, the orchestrator increments a retry counter stored in the task metadata. The system allows a maximum of 3 attempts per task. Each retry iteration includes the previous QA feedback—specific issues, screenshots, logs, and fix instructions—ensuring the developer agent has full context for the next attempt.
This logic is defined in specialized/agents-orchestrator.md under Failure Management:
# agents-orchestrator.md – Failure handling excerpt
failure_handling:
max_retries: 3 # ≤ 3 attempts per task
on_fail:
- increment_retry_counter
- if_retry < max_retries: # loop back to developer with feedback
route_to: developer_agent
include: qa_feedback
- else: # 3rd failure → escalation
generate: escalation_report
route_to: studio_producer
Automatic Escalation Triggers
After the third failed attempt (retry >= 3), the orchestrator automatically triggers the Escalation Protocol. At this stage, the task is no longer routed back to the original developer. Instead, the system generates a structured Escalation Report using the markdown template defined in strategy/nexus-strategy.md and forwards it to the Studio Producer or another designated gate-keeper.
Escalation Protocol and Gate-Keeper Intervention
Escalation Report Generation
The Escalation Report is a comprehensive markdown document that captures the full failure history, root-cause analysis, and recommended resolutions. According to strategy/nexus-strategy.md, the report must include:
- Failure History: Chronological log of all three attempts with QA feedback
- Root Cause Analysis: Technical assessment of why the fix failed
- Recommended Resolutions: Specific actions such as re-assignment, task decomposition, or architecture revision
- Impact Assessment: Timeline and quality metrics affected by the blockage
Gate-Keeper Decision Matrix
Once the Studio Producer receives the Escalation Report, they evaluate the issue using the decision matrix defined in the Escalation Protocol section of strategy/nexus-strategy.md. The gate-keeper has four primary resolution paths:
- Re-assign: Route the task to a different developer agent with specialized skills
- Decompose: Split the complex task into smaller sub-tasks that can pass QA independently
- Revise Architecture: Engage architect agents to adjust technical specifications or infrastructure
- Accept Limitation: Mark the task as BLOCKED with documented constraints and move forward
The support/support-support-responder.md file provides additional criteria for when to involve senior leadership or external resources, ensuring that business-critical blockages receive appropriate visibility.
Configuration and Implementation Details
QA Failure Feedback Template
When QA returns a failure, the feedback follows a strict template defined in strategy/nexus-strategy.md to ensure the developer receives actionable data:
## QA Failure Feedback
### Task: [Task ID – e.g., FE-42]
### Attempt: 2/3
### Verdict: FAIL
#### Specific Issues Found
1. **Layout Shift**: Header overlaps content on mobile.
- Expected: Fixed header stays above content.
- Actual: Header pushes content down.
- Evidence: `header_shift.png`
2. **API Timeout**: `/api/v1/items` returns 504 after 30 s.
- Expected: ≤ 200 ms response.
- Actual: 30 s timeout.
- Evidence: `api_timeout.log`
#### Fix Instructions
- Adjust CSS `position: sticky` on header.
- Optimize DB query for `/api/v1/items` (add index on `status`).
#### Files to Modify
- `src/components/Header.vue`
- `src/api/items.js`
#### Retry Expectations
- Implement fixes only; do **not** add new features.
- Submit for QA as attempt **3/3`.
Escalation Report Template
The final escalation document used by the Studio Producer follows this structure from strategy/nexus-strategy.md:
## Escalation Report
### Task: FE-42 – Mobile Header Layout
### Attempts Exhausted: 3/3
### Escalation Level: To Studio Producer
#### Failure History
- **Attempt 1**: Header overlaps; no visual proof.
- **Attempt 2**: Same issue; added screenshots.
- **Attempt 3**: Still fails after CSS refactor.
#### Root Cause Analysis
- CSS `position: sticky` not compatible with current layout grid.
- Backend API latency due to missing index.
#### Recommended Resolution
- [ ] Reassign to a Senior Frontend Developer.
- [ ] Decompose UI fix into two sub‑tasks (header, grid).
- [ ] backend team adds index on `status`.
- [ ] Schedule a sprint‑level design review.
#### Impact Assessment
- **Blocking**: Subsequent UI screens depend on fixed header.
- **Timeline Impact**: +2 days to sprint.
- **Quality Impact**: Current QA pass rate drops to 68 %.
Summary
- The Agents Orchestrator enforces a 3-attempt limit per task, cycling failed work through a Dev↔QA feedback loop defined in
specialized/agents-orchestrator.md. - QA failures trigger detailed feedback reports that include specific issues, evidence files, and fix instructions to guide the next retry attempt.
- After the third failure, the system automatically generates an Escalation Report and routes the task to a gate-keeper such as the Studio Producer.
- The gate-keeper can re-assign, decompose, revise architecture, or accept limitations based on the escalation protocol in
strategy/nexus-strategy.md. - Support Responders provide additional routing criteria for involving senior leadership or external resources when business-critical blockages occur.
Frequently Asked Questions
What happens when a task fails QA three times?
After the third QA failure, the Agents Orchestrator stops the retry loop and triggers the Escalation Protocol. The system generates a structured Escalation Report containing the full failure history, root-cause analysis, and recommended resolutions, then routes the task to the Studio Producer or another designated gate-keeper for manual intervention.
Who receives the Escalation Report and what can they do?
The Studio Producer typically receives the Escalation Report, though the routing may vary based on criteria in support/support-support-responder.md. The gate-keeper has four primary resolution options: re-assign the task to a different developer agent, decompose the task into smaller sub-tasks, revise the technical architecture through architect agents, or accept the limitation and mark the task as blocked.
How does the QA feedback loop improve failure recovery?
Each retry attempt includes a detailed QA Failure Feedback template that specifies exact issues found, evidence files (screenshots, logs), expected vs. actual behavior, and precise fix instructions. This ensures the developer agent receives actionable context rather than generic error messages, maximizing the probability of success on subsequent attempts within the three-retry limit.
Can retry limits be configured for different task types?
The default configuration in specialized/agents-orchestrator.md sets a universal maximum of 3 retries per task before mandatory escalation. While the source documentation emphasizes this consistent limit to maintain pipeline velocity, specific task types or critical paths could theoretically implement custom retry logic by modifying the failure_handling configuration block in the orchestrator's definition file.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →