NVIDIA Plugin Orchestrator Skill Pattern: Phase-Split Workflow Architecture
The NVIDIA plugin implements a deterministic, self-healing Orchestrator skill pattern that separates workflow submission, monitoring, and failure diagnosis into distinct phases across specialized sub-agents.
The Orchestrator skill pattern demonstrated in the openai/plugins repository provides a robust framework for autonomous workflow management within the NVIDIA plugin. This architecture isolates mutation operations to dedicated agents while maintaining continuous observability through the main skill. By splitting responsibilities across the workflow-expert and logs-reader sub-agents, the pattern enables resilient execution of OSMO workflows with automatic failure recovery.
Phase-Split Architecture Overview
The Orchestrator skill pattern operates through four distinct phases, each handled by specific agents defined in SKILL.md and the agent definitions under components/osmo-cli/agents/. This separation ensures that submission logic remains decoupled from monitoring and diagnostic responsibilities.
Phase 1: Workflow Submission with workflow-expert
The workflow-expert sub-agent, defined in components/osmo-cli/agents/workflow-expert.md, handles the setup and submission of OSMO workflows. Crucially, this agent does not monitor the execution; it validates resources, submits the workflow using osmo workflow submit, and returns the workflow ID, pool name, and resumption context to the main skill.
# components/osmo-cli/agents/workflow-expert.md
steps:
- name: preflight
command: osmo preflight --check
- name: submit
command: osmo workflow submit workflow.yaml --pool $POOL
capture: workflow_id
- name: output
echo: "Submitted workflow $workflow_id"
Phase 2: Inline Monitoring by Main Skill
While sub-agents execute, the main skill (physical-ai-infrastructure-setup-and-resilient-scaling) performs inline monitoring by polling osmo workflow query <id> and streaming status updates to the user. This design guarantees live progress visibility rather than hiding activity in background processes.
# Main skill monitors the workflow inline
while true; do
osmo workflow query $WORKFLOW_ID --format-type json | jq .
sleep 5
done
Phase 3: Failure Diagnosis with logs-reader
When workflows fail, the Orchestrator either resumes the workflow-expert or launches the logs-reader sub-agent to perform root cause analysis. Located at components/osmo-cli/agents/logs-reader.md, this agent fetches logs using osmo workflow logs, identifies specific failure modes (such as missing jq commands or incorrect {{outputs}} placeholders), and edits the workflow.yaml file directly.
# components/osmo-cli/agents/logs-reader.md
steps:
- name: fetch-logs
command: osmo workflow logs $WORKFLOW_ID -n 10000 > logs.txt
- name: diagnose
script: |
if grep -q "jq: command not found" logs.txt; then
echo "Install jq or replace its usage"
fi
if grep -q "{{outputs}}" logs.txt; then
echo "Replace {{outputs}} with {{output}}"
fi
- name: fix-workflow
edit: workflow.yaml
Phase 4: Self-Healing Retry Loop
The main skill implements a deterministic retry loop that resubmits corrected workflows until completion or retry exhaustion. After the logs-reader applies fixes (such as installing missing tools or correcting placeholder syntax), the main skill resumes the workflow-expert with the failure context and re-submits the workflow, creating a closed-loop healing mechanism.
Implementation Files and References
According to the source code in the openai/plugins repository, the Orchestrator skill pattern is defined across several key files:
plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/SKILL.md– Contains the top-level skill definition and the Component References table listing the sub-agents.plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/agents/workflow-expert.md– Implements the submission phase.plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/agents/logs-reader.md– Implements the diagnostic and repair phase.plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/tests/orchestrator-runtime-failure.md– Validates the complete end-to-end orchestration flow.
Practical Implementation Example
The following test case from orchestrator-runtime-failure.md demonstrates the complete Orchestrator skill pattern in action:
# Test prompt driving the orchestrator pattern
prompt: |
I have a hello world workflow ready in workflow.yaml. Submit it now,
monitor it, and if it fails, diagnose, fix, and resubmit automatically.
This prompt triggers the main skill to:
- Spawn the
workflow-expertto submit the workflow - Poll status every 5 seconds via
osmo workflow query - Invoke
logs-readerupon failure to analyze logs and patchworkflow.yaml - Retry submission until success or limit reached
Summary
The Orchestrator skill pattern in the NVIDIA plugin provides a deterministic framework for resilient workflow management:
- Phase separation isolates submission (
workflow-expert), monitoring (main skill), and diagnosis (logs-reader) into distinct responsibilities. - Inline polling ensures users receive real-time status updates rather than opaque background processing.
- Self-healing loops automatically diagnose failures (missing dependencies, placeholder errors) and re-submit corrected workflows.
- Deterministic architecture guarantees reproducible behavior across retries, with clear mutation boundaries between agents.
Frequently Asked Questions
What makes the Orchestrator skill pattern different from standard agent delegation?
Unlike simple task delegation, the Orchestrator skill pattern explicitly separates mutation (workflow submission and editing) from observability (status polling). The main skill never directly modifies workflow files; instead, it orchestrates specialized sub-agents (workflow-expert for submission, logs-reader for fixes) while maintaining control of the monitoring loop. This creates cleaner error boundaries and enables automatic retry logic without state pollution.
How does the main skill handle workflow failures without losing context?
When a failure is detected through osmo workflow query, the main skill captures the workflow ID and error state, then either resumes the existing workflow-expert instance or initializes the logs-reader agent with the specific failure context. The sub-agents receive the full execution history and file paths, allowing them to perform surgical edits to workflow.yaml (such as replacing {{outputs}} with {{output}}) before the main skill triggers re-submission.
Where is the Orchestrator pattern tested in the NVIDIA plugin?
The pattern is validated in components/osmo-cli/tests/orchestrator-runtime-failure.md, which tests the complete lifecycle: initial submission through workflow-expert, simulated runtime failure, automated diagnosis via logs-reader, and successful retry. This test ensures that the phase-split behavior correctly handles real-world failure modes like missing system dependencies or template syntax errors.
Can the Orchestrator pattern handle multiple workflow failures in sequence?
Yes, the implementation includes a retry limit that loops through phases 3 and 4 repeatedly. For each failure, the logs-reader performs incremental fixes (documented in its diagnostic script), and the main skill re-submits via the workflow-expert. This loop continues until the workflow completes successfully or exhausts the configured retry threshold, making the system resilient against cascading or compound failures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →