How Oh-My-Codex Executes Parallel Agent Delegation in Ralph Mode Using Agent Tiers

Ralph mode executes parallel agent delegation by generating a staffing plan in src/team/followup-planner.ts that assigns agents to LOW, STANDARD, or THOROUGH reasoning-effort tiers, then spawning simultaneous Codex native subagents via the <ralph_native_subagents> instruction block while coordinating completion through a verification lane.

The open-source oh-my-codex project provides a command-line interface for orchestrating multiple AI agents through a persistence mode called Ralph. When you invoke omx ralph, the system automatically distributes tasks across specialized agents using a tier-based parallelism strategy defined in the source code. This implementation enables complex software engineering tasks to run concurrently rather than sequentially, significantly reducing total execution time.


# 1️⃣ Start Ralph with a task – the CLI builds the staffing plan and writes the subagent block

omx ralph "Add OAuth2 support to auth module"

Staffing Plan Generation and Role Allocation

The parallel agent delegation process begins in src/team/followup-planner.ts with the buildFollowupStaffingPlan() function. This utility inspects available prompt-defined agents and constructs an allocations array that designates specific roles—such as a primary implementation role and a quality verification role—along with worker counts. Each allocation entry includes a reasoning effort tier derived from the agent's reasoningEffort field defined in src/agents/definitions.ts, categorizing the workload as LOW, STANDARD, or THOROUGH.

// src/team/followup-planner.ts – build staffing plan (excerpt)
export function buildFollowupStaffingPlan(
  mode: FollowupMode,
  task: string,
  availableAgentTypes: readonly string[],
  options: BuildFollowupStaffingPlanOptions = {},
): FollowupStaffingPlan {
  const workerCount = Math.max(1, options.workerCount ?? (mode === 'team' ? 2 : 3));
  const primaryRole = chooseAvailableRole(availableAgentTypes, [primaryRoute.role], fallbackRole);
  const qualityRole = chooseAvailableRole(availableAgentTypes, ['test-engineer','verifier','quality-reviewer'], primaryRole);
  // … allocate workers with appropriate tier (reasoningEffort) …
}

For example, an allocation object for an executor role appears as { role: "executor", count: 1, reason: "primary implementation lane", reasoningEffort: "high" }, which maps to the THOROUGH tier. The function also accepts a workerCount option to scale parallelism, defaulting to three workers in Ralph mode.

Native Subagent Enablement and Session Configuration

Once the staffing plan is finalized, src/cli/ralph.ts prepares the runtime environment for parallel execution. The writeRalphSessionFiles() function generates a session instruction file at .omx/ralph/session-instructions.md containing a Ralph native-subagents block wrapped in <ralph_native_subagents> tags. This block embeds task context, parallelism guidance, and references to the subagent activity ledger at .omx/state/subagent-tracking.json.

// src/cli/ralph.ts – enable native subagents and write the instruction block
await updateModeState('ralph', {
  …,
  native_subagents_enabled: true,                     // ← enables parallel subagents
  native_subagent_tracking_path: '.omx/state/subagent-tracking.json',
});
const sessionFiles = await writeRalphSessionFiles(cwd, task, { noDeslop, approvedHint });
process.env[RALPH_APPEND_ENV] = sessionFiles.instructionsPath; // inject block
await launchWithHud(codexArgs);

The CLI then invokes updateModeState() to set native_subagents_enabled: true (line 190) and exports the session file path via the OMX_RALPH_APPEND_INSTRUCTIONS_FILE environment variable. This configuration instructs the Codex executor to parse the subagent block and spawn native background workers.


# Inside the generated `.omx/ralph/session-instructions.md`

<ralph_native_subagents>
You are in OMX Ralph persistence mode.
Primary task: Add OAuth2 support to auth module
Parallelism guidance:
- Prefer Codex native subagents for independent parallel subtasks.
- Treat `.omx/state/subagent-tracking.json` as the native subagent activity ledger for this session.
…
</ralph_native_subagents>

Tier-Driven Model Selection

The oh-my-codex repository maps each agent tier to a specific model class through src/utils/agents-model-table.ts. The getAgentRecommendedModel() function selects the appropriate model based on the agent's modelClass property, which aligns with the three reasoning-effort tiers established in the staffing plan.

LOW Tier: Fast Models

Agents with reasoningEffort: "low", such as exploration-focused roles, are assigned to the LOW tier and utilize the sparkModel class. This configuration prioritizes speed for quick look-ups and preliminary research tasks.

STANDARD Tier: Default Models

Agents carrying reasoningEffort: "medium", including debugger and quality-reviewer roles, fall into the STANDARD tier. These workers run on the subagentDefaultModel class, balancing cost and capability for general implementation and verification tasks.

THOROUGH Tier: Frontier Models

Complex architectural and implementation tasks assigned reasoningEffort: "high"—such as architect and executor roles—map to the THOROUGH tier. These agents receive the frontierModel class, deploying the most capable available models for deep reasoning and critical code generation.

Parallel Execution and Verification Coordination

With tier-specific models selected, the Codex runtime launches each subagent simultaneously using run_in_background: true as defined in the skill configuration (skills/ralph/SKILL.md). Independent lanes for implementation, verification, and specialist support execute concurrently rather than blocking on sequential completion.

According to the buildVerificationPlan() logic in src/team/followup-planner.ts, the Ralph persistence loop monitors only the verification lane for completion before proceeding. Once verification evidence passes and all parallel work converges, the system issues a /cancel command to clean up state and finalize the session.

Summary

  • Ralph mode activates parallel delegation by setting native_subagents_enabled: true in src/cli/ralph.ts and exporting session instructions via OMX_RALPH_APPEND_INSTRUCTIONS_FILE.
  • The buildFollowupStaffingPlan() function in src/team/followup-planner.ts creates tier-aware allocations based on agent reasoningEffort values defined in src/agents/definitions.ts.
  • Three agent tiers—LOW, STANDARD, and THOROUGH—map to sparkModel, subagentDefaultModel, and frontierModel respectively via src/utils/agents-model-table.ts.
  • Subagents launch simultaneously using background execution, with the Ralph loop waiting specifically for the verification lane to complete before finalizing.

Frequently Asked Questions

How does Ralph mode decide which agents to run in parallel?

The buildFollowupStaffingPlan() function in src/team/followup-planner.ts analyzes the available agent types and task requirements to select a primary implementation role and a quality verification role. It returns an allocations array containing each role, worker count, and reasoning-effort tier, which determines how many parallel subagents to spawn.

What are the three agent tiers in oh-my-codex and how do they differ?

The three tiers are LOW (fast sparkModel for exploration), STANDARD (balanced subagentDefaultModel for debugging and review), and THOROUGH (capable frontierModel for architecture and execution). Each tier corresponds to the reasoningEffort field—"low", "medium", or "high"—defined in src/agents/definitions.ts and selects appropriate compute resources.

Where does Ralph mode store subagent tracking information?

Ralph writes subagent activity data to .omx/state/subagent-tracking.json, referenced as native_subagent_tracking_path in the mode state. The session instructions at .omx/ralph/session-instructions.md direct subagents to treat this file as the activity ledger for the current session.

Does Ralph wait for all subagents to finish before continuing?

No, the Ralph loop specifically waits only for the verification lane to complete, as implemented in the verification plan logic. Implementation and specialist lanes may still be running in the background, but the primary thread proceeds once quality checks are satisfied, optimizing total execution time.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →