# NVIDIA Plugin Orchestrator Skill Pattern: Phase-Split Workflow Architecture

> Explore the NVIDIA plugin's Orchestrator skill pattern. Learn how phase-split architecture separates workflow submission, monitoring, and failure diagnosis for self-healing.

- Repository: [OpenAI/plugins](https://github.com/openai/plugins)
- Tags: architecture
- Published: 2026-09-10

---

**The NVIDIA plugin implements a deterministic, self-healing Orchestrator skill pattern that separates workflow submission, monitoring, and failure diagnosis into distinct phases across specialized sub-agents.**

The **Orchestrator skill pattern** demonstrated in the openai/plugins repository provides a robust framework for autonomous workflow management within the NVIDIA plugin. This architecture isolates mutation operations to dedicated agents while maintaining continuous observability through the main skill. By splitting responsibilities across the `workflow-expert` and `logs-reader` sub-agents, the pattern enables resilient execution of OSMO workflows with automatic failure recovery.

## Phase-Split Architecture Overview

The **Orchestrator skill pattern** operates through four distinct phases, each handled by specific agents defined in [`SKILL.md`](https://github.com/openai/plugins/blob/main/SKILL.md) and the agent definitions under `components/osmo-cli/agents/`. This separation ensures that submission logic remains decoupled from monitoring and diagnostic responsibilities.

### Phase 1: Workflow Submission with workflow-expert

The `workflow-expert` sub-agent, defined in [`components/osmo-cli/agents/workflow-expert.md`](https://github.com/openai/plugins/blob/main/components/osmo-cli/agents/workflow-expert.md), handles the **setup** and **submission** of OSMO workflows. Crucially, this agent does not monitor the execution; it validates resources, submits the workflow using `osmo workflow submit`, and returns the workflow ID, pool name, and resumption context to the main skill.

```yaml

# components/osmo-cli/agents/workflow-expert.md

steps:
  - name: preflight
    command: osmo preflight --check
  - name: submit
    command: osmo workflow submit workflow.yaml --pool $POOL
    capture: workflow_id
  - name: output
    echo: "Submitted workflow $workflow_id"

```

### Phase 2: Inline Monitoring by Main Skill

While sub-agents execute, the main skill (`physical-ai-infrastructure-setup-and-resilient-scaling`) performs **inline monitoring** by polling `osmo workflow query <id>` and streaming status updates to the user. This design guarantees live progress visibility rather than hiding activity in background processes.

```bash

# Main skill monitors the workflow inline

while true; do
  osmo workflow query $WORKFLOW_ID --format-type json | jq .
  sleep 5
done

```

### Phase 3: Failure Diagnosis with logs-reader

When workflows fail, the **Orchestrator** either resumes the `workflow-expert` or launches the `logs-reader` sub-agent to perform root cause analysis. Located at [`components/osmo-cli/agents/logs-reader.md`](https://github.com/openai/plugins/blob/main/components/osmo-cli/agents/logs-reader.md), this agent fetches logs using `osmo workflow logs`, identifies specific failure modes (such as missing `jq` commands or incorrect `{{outputs}}` placeholders), and edits the [`workflow.yaml`](https://github.com/openai/plugins/blob/main/workflow.yaml) file directly.

```yaml

# components/osmo-cli/agents/logs-reader.md

steps:
  - name: fetch-logs
    command: osmo workflow logs $WORKFLOW_ID -n 10000 > logs.txt
  - name: diagnose
    script: |
      if grep -q "jq: command not found" logs.txt; then
        echo "Install jq or replace its usage"
      fi
      if grep -q "{{outputs}}" logs.txt; then
        echo "Replace {{outputs}} with {{output}}"
      fi
  - name: fix-workflow
    edit: workflow.yaml

```

### Phase 4: Self-Healing Retry Loop

The main skill implements a **deterministic retry loop** that resubmits corrected workflows until completion or retry exhaustion. After the `logs-reader` applies fixes (such as installing missing tools or correcting placeholder syntax), the main skill resumes the `workflow-expert` with the failure context and re-submits the workflow, creating a closed-loop healing mechanism.

## Implementation Files and References

According to the source code in the openai/plugins repository, the **Orchestrator skill pattern** is defined across several key files:

- **[`plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/SKILL.md`](https://github.com/openai/plugins/blob/main/plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/SKILL.md)** – Contains the top-level skill definition and the *Component References* table listing the sub-agents.
- **[`plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/agents/workflow-expert.md`](https://github.com/openai/plugins/blob/main/plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/agents/workflow-expert.md)** – Implements the submission phase.
- **[`plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/agents/logs-reader.md`](https://github.com/openai/plugins/blob/main/plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/agents/logs-reader.md)** – Implements the diagnostic and repair phase.
- **[`plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/tests/orchestrator-runtime-failure.md`](https://github.com/openai/plugins/blob/main/plugins/nvidia/skills/physical-ai-infrastructure-setup-and-resilient-scaling/components/osmo-cli/tests/orchestrator-runtime-failure.md)** – Validates the complete end-to-end orchestration flow.

## Practical Implementation Example

The following test case from [`orchestrator-runtime-failure.md`](https://github.com/openai/plugins/blob/main/orchestrator-runtime-failure.md) demonstrates the complete **Orchestrator skill pattern** in action:

```yaml

# Test prompt driving the orchestrator pattern

prompt: |
  I have a hello world workflow ready in workflow.yaml. Submit it now,
  monitor it, and if it fails, diagnose, fix, and resubmit automatically.

```

This prompt triggers the main skill to:
1. Spawn the `workflow-expert` to submit the workflow
2. Poll status every 5 seconds via `osmo workflow query`
3. Invoke `logs-reader` upon failure to analyze logs and patch [`workflow.yaml`](https://github.com/openai/plugins/blob/main/workflow.yaml)
4. Retry submission until success or limit reached

## Summary

The **Orchestrator skill pattern** in the NVIDIA plugin provides a deterministic framework for resilient workflow management:

- **Phase separation** isolates submission (`workflow-expert`), monitoring (main skill), and diagnosis (`logs-reader`) into distinct responsibilities.
- **Inline polling** ensures users receive real-time status updates rather than opaque background processing.
- **Self-healing loops** automatically diagnose failures (missing dependencies, placeholder errors) and re-submit corrected workflows.
- **Deterministic architecture** guarantees reproducible behavior across retries, with clear mutation boundaries between agents.

## Frequently Asked Questions

### What makes the Orchestrator skill pattern different from standard agent delegation?

Unlike simple task delegation, the **Orchestrator skill pattern** explicitly separates **mutation** (workflow submission and editing) from **observability** (status polling). The main skill never directly modifies workflow files; instead, it orchestrates specialized sub-agents (`workflow-expert` for submission, `logs-reader` for fixes) while maintaining control of the monitoring loop. This creates cleaner error boundaries and enables automatic retry logic without state pollution.

### How does the main skill handle workflow failures without losing context?

When a failure is detected through `osmo workflow query`, the main skill captures the workflow ID and error state, then either resumes the existing `workflow-expert` instance or initializes the `logs-reader` agent with the specific failure context. The sub-agents receive the full execution history and file paths, allowing them to perform surgical edits to [`workflow.yaml`](https://github.com/openai/plugins/blob/main/workflow.yaml) (such as replacing `{{outputs}}` with `{{output}}`) before the main skill triggers re-submission.

### Where is the Orchestrator pattern tested in the NVIDIA plugin?

The pattern is validated in [`components/osmo-cli/tests/orchestrator-runtime-failure.md`](https://github.com/openai/plugins/blob/main/components/osmo-cli/tests/orchestrator-runtime-failure.md), which tests the complete lifecycle: initial submission through `workflow-expert`, simulated runtime failure, automated diagnosis via `logs-reader`, and successful retry. This test ensures that the phase-split behavior correctly handles real-world failure modes like missing system dependencies or template syntax errors.

### Can the Orchestrator pattern handle multiple workflow failures in sequence?

Yes, the implementation includes a retry limit that loops through phases 3 and 4 repeatedly. For each failure, the `logs-reader` performs incremental fixes (documented in its diagnostic script), and the main skill re-submits via the `workflow-expert`. This loop continues until the workflow completes successfully or exhausts the configured retry threshold, making the system resilient against cascading or compound failures.