# How Ouroboros Parallel Execution Handles Multiple LLM Models Simultaneously

> Learn how Ouroboros Parallel Execution simultaneously manages multiple LLM models using PAL router and isolated async sessions. Discover efficient LLM orchestration.

- Repository: [Q00/ouroboros](https://github.com/Q00/ouroboros)
- Tags: deep-dive
- Published: 2026-03-14

---

**Ouroboros Parallel Execution orchestrates multiple LLM models simultaneously by routing each Acceptance Criterion to a model-specific tier through the PAL router, then executing them in isolated async sessions coordinated by the ParallelACExecutor and LevelCoordinator.**

The Q00/ouroboros repository implements a sophisticated parallel execution architecture that enables simultaneous processing across heterogeneous LLM endpoints. By combining dependency-aware scheduling with intelligent model routing, Ouroboros Parallel Execution distributes Acceptance Criteria across different model tiers—from Frugal to Frontier—within a single parallel batch. This design ensures optimal resource utilization while maintaining deterministic orchestration and conflict resolution.

## Dependency-Aware Scheduling for Parallel Execution

Parallel execution begins with the **DependencyAnalyzer** class located in [`src/ouroboros/orchestrator/dependency_analyzer.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/orchestrator/dependency_analyzer.py). This component sends every Acceptance Criterion (AC) to a Claude LLM to generate a JSON dependency map, which it then converts into a **DependencyGraph** data structure.

Each level of the graph contains ACs that can run without waiting for one another. The **ParallelACExecutor** processes these levels sequentially, but executes all ACs within a single level simultaneously. This dependency-aware approach ensures that parallel execution only occurs where semantically safe, preventing race conditions on interdependent tasks while maximizing throughput for independent ones.

## Per-AC Model Routing with the PAL Router

Before an AC is handed to the execution adapter, the **PAL router** ([`src/ouroboros/routing/router.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/routing/router.py)) evaluates specific task characteristics to determine the optimal LLM endpoint.

### Evaluating Task Complexity

The router analyzes **task complexity** across three dimensions:

- **Token count** of the input context
- **Tool dependencies** required for execution
- **AC depth** within the decomposition hierarchy

Based on this analysis, the system assigns each AC to one of three model tiers: **Frugal**, **Standard**, or **Frontier**.

### Resolving Model Tiers to Concrete Endpoints

The abstract tiers defined in the routing logic resolve to concrete model configurations in [`src/ouroboros/config/models.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/config/models.py). For example:

- **Frugal** tier maps to `gpt-4o-mini`
- **Standard** tier maps to `claude-sonnet-4-6`
- **Frontier** tier maps to `gemini-2.5-pro`

This resolution happens independently for each AC, meaning a single parallel batch may simultaneously invoke OpenAI, Anthropic, and Google models according to their specific requirements.

## Concurrent Multi-Model Execution

Once routed, the **ParallelACExecutor** ([`src/ouroboros/orchestrator/parallel_executor.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/orchestrator/parallel_executor.py)) manages the actual concurrent processing across these heterogeneous endpoints.

### Isolated Provider Sessions

Within each execution level, the executor spawns an async task for every AC. Each task calls `self._adapter.execute_task(...)`, which creates a **separate Claude session** (or other provider-specific session) for that specific AC (see lines 27-45 in [`src/ouroboros/orchestrator/parallel_executor.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/orchestrator/parallel_executor.py)). Because the adapter receives the model name dynamically from the router, different ACs in the same parallel batch process through different LLM endpoints without session interference or resource contention.

### Async Task Orchestration

The parallel executor groups ACs by dependency level, then uses Python's `asyncio` infrastructure to launch all ACs in the current level concurrently. Each maintains its own connection pool and authentication context to its assigned provider, allowing true parallel execution across multiple LLM APIs rather than sequential round-robin processing.

## Conflict Resolution via the Coordinator Agent

After a level completes execution, the **LevelCoordinator** scans the message history for file-modifying tool calls. If two ACs modified the same file, the system launches a single **Coordinator Claude session** to resolve the conflict.

This mediation step does not affect the already-running AC sessions—they continue processing in parallel with whatever model they were assigned. The coordinator operates as a separate reconciliation layer, merging conflicting changes before the next dependency level begins execution.

## Practical Implementation Examples

The following examples demonstrate how to configure multi-model parallel execution in Ouroboros:

```python

# example seed (parallel-ready) – different ACs will be routed to different models

seed_yaml = """
task_type: code
acceptance_criteria:
  - "AC 1: Create a FastAPI entrypoint (use Claude-sonnet-4-6)"
  - "AC 2: Add a Postgres migration script (use gpt-4o-mini)"
  - "AC 3: Write unit tests for the endpoint (use gemini-2.5-pro)"
"""

# CLI invocation – parallel execution enabled (default)

# The router will auto-select the model for each AC

!ouroboros run workflow --orchestrator -f seed.yaml

```

```python

# programmatic use – manually set a router and executor

from ouroboros.routing import PALRouter
from ouroboros.orchestrator.runner import OrchestratorRunner
from ouroboros.orchestrator.adapter import ClaudeAgentAdapter

router = PALRouter()

# resolve a model for a dummy task (shows which model would be used)

decision = router.route(task_context)           # → RoutingDecision(tier=Tier.STANDARD, …)

adapter = ClaudeAgentAdapter()                 # uses the model supplied by the router

runner = OrchestratorRunner(adapter, event_store)
await runner.execute_seed(seed, parallel=True)  # each AC runs in its own Claude session

```

## Summary

- **DependencyAnalyzer** in [`src/ouroboros/orchestrator/dependency_analyzer.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/orchestrator/dependency_analyzer.py) generates execution levels where ACs within each level run in parallel.
- **PALRouter** in [`src/ouroboros/routing/router.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/routing/router.py) evaluates task complexity and assigns each AC to a specific model tier (Frugal, Standard, or Frontier).
- **ParallelACExecutor** spawns isolated async tasks for each AC, creating separate provider sessions that allow different models to run simultaneously.
- **LevelCoordinator** resolves file conflicts post-execution without interrupting the parallel processing of other ACs.
- Model configurations in [`src/ouroboros/config/models.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/config/models.py) map abstract tiers to concrete endpoints like `gpt-4o-mini`, `claude-sonnet-4-6`, and `gemini-2.5-pro`.

## Frequently Asked Questions

### How does Ouroboros decide which LLM model to use for each Acceptance Criterion?

The **PALRouter** analyzes task complexity based on token count, tool dependencies, and AC depth, then assigns a tier (Frugal, Standard, or Frontier). This tier resolves to a specific model defined in [`src/ouroboros/config/models.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/config/models.py), such as routing lightweight tasks to `gpt-4o-mini` and complex reasoning tasks to `gemini-2.5-pro`.

### Can different ACs in the same parallel batch use different model providers?

Yes. Because the **ClaudeAgentAdapter** receives the model name dynamically from the router and creates a separate session for each AC, a single parallel batch can simultaneously execute ACs through OpenAI, Anthropic, and Google endpoints. Each maintains its own isolated connection without cross-provider interference.

### What happens if two parallel ACs modify the same file?

The **LevelCoordinator** detects file-modifying tool call conflicts after a level completes. It launches a dedicated **Coordinator Claude session** to reconcile the changes before proceeding to the next dependency level. This resolution occurs independently and does not block or restart the other parallel AC executions.

### How does the PAL router determine the appropriate model tier?

The router evaluates three specific metrics: the **token count** of the task context, the **tool dependencies** required (such as file system or database access), and the **AC depth** within the task decomposition hierarchy. These factors collectively determine whether an AC receives Frugal, Standard, or Frontier tier allocation according to the logic in [`src/ouroboros/routing/router.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/routing/router.py).