How Ouroboros Parallel Execution Handles Multiple LLM Models Simultaneously
Ouroboros Parallel Execution orchestrates multiple LLM models simultaneously by routing each Acceptance Criterion to a model-specific tier through the PAL router, then executing them in isolated async sessions coordinated by the ParallelACExecutor and LevelCoordinator.
The Q00/ouroboros repository implements a sophisticated parallel execution architecture that enables simultaneous processing across heterogeneous LLM endpoints. By combining dependency-aware scheduling with intelligent model routing, Ouroboros Parallel Execution distributes Acceptance Criteria across different model tiers—from Frugal to Frontier—within a single parallel batch. This design ensures optimal resource utilization while maintaining deterministic orchestration and conflict resolution.
Dependency-Aware Scheduling for Parallel Execution
Parallel execution begins with the DependencyAnalyzer class located in src/ouroboros/orchestrator/dependency_analyzer.py. This component sends every Acceptance Criterion (AC) to a Claude LLM to generate a JSON dependency map, which it then converts into a DependencyGraph data structure.
Each level of the graph contains ACs that can run without waiting for one another. The ParallelACExecutor processes these levels sequentially, but executes all ACs within a single level simultaneously. This dependency-aware approach ensures that parallel execution only occurs where semantically safe, preventing race conditions on interdependent tasks while maximizing throughput for independent ones.
Per-AC Model Routing with the PAL Router
Before an AC is handed to the execution adapter, the PAL router (src/ouroboros/routing/router.py) evaluates specific task characteristics to determine the optimal LLM endpoint.
Evaluating Task Complexity
The router analyzes task complexity across three dimensions:
- Token count of the input context
- Tool dependencies required for execution
- AC depth within the decomposition hierarchy
Based on this analysis, the system assigns each AC to one of three model tiers: Frugal, Standard, or Frontier.
Resolving Model Tiers to Concrete Endpoints
The abstract tiers defined in the routing logic resolve to concrete model configurations in src/ouroboros/config/models.py. For example:
- Frugal tier maps to
gpt-4o-mini - Standard tier maps to
claude-sonnet-4-6 - Frontier tier maps to
gemini-2.5-pro
This resolution happens independently for each AC, meaning a single parallel batch may simultaneously invoke OpenAI, Anthropic, and Google models according to their specific requirements.
Concurrent Multi-Model Execution
Once routed, the ParallelACExecutor (src/ouroboros/orchestrator/parallel_executor.py) manages the actual concurrent processing across these heterogeneous endpoints.
Isolated Provider Sessions
Within each execution level, the executor spawns an async task for every AC. Each task calls self._adapter.execute_task(...), which creates a separate Claude session (or other provider-specific session) for that specific AC (see lines 27-45 in src/ouroboros/orchestrator/parallel_executor.py). Because the adapter receives the model name dynamically from the router, different ACs in the same parallel batch process through different LLM endpoints without session interference or resource contention.
Async Task Orchestration
The parallel executor groups ACs by dependency level, then uses Python's asyncio infrastructure to launch all ACs in the current level concurrently. Each maintains its own connection pool and authentication context to its assigned provider, allowing true parallel execution across multiple LLM APIs rather than sequential round-robin processing.
Conflict Resolution via the Coordinator Agent
After a level completes execution, the LevelCoordinator scans the message history for file-modifying tool calls. If two ACs modified the same file, the system launches a single Coordinator Claude session to resolve the conflict.
This mediation step does not affect the already-running AC sessions—they continue processing in parallel with whatever model they were assigned. The coordinator operates as a separate reconciliation layer, merging conflicting changes before the next dependency level begins execution.
Practical Implementation Examples
The following examples demonstrate how to configure multi-model parallel execution in Ouroboros:
# example seed (parallel-ready) – different ACs will be routed to different models
seed_yaml = """
task_type: code
acceptance_criteria:
- "AC 1: Create a FastAPI entrypoint (use Claude-sonnet-4-6)"
- "AC 2: Add a Postgres migration script (use gpt-4o-mini)"
- "AC 3: Write unit tests for the endpoint (use gemini-2.5-pro)"
"""
# CLI invocation – parallel execution enabled (default)
# The router will auto-select the model for each AC
!ouroboros run workflow --orchestrator -f seed.yaml
# programmatic use – manually set a router and executor
from ouroboros.routing import PALRouter
from ouroboros.orchestrator.runner import OrchestratorRunner
from ouroboros.orchestrator.adapter import ClaudeAgentAdapter
router = PALRouter()
# resolve a model for a dummy task (shows which model would be used)
decision = router.route(task_context) # → RoutingDecision(tier=Tier.STANDARD, …)
adapter = ClaudeAgentAdapter() # uses the model supplied by the router
runner = OrchestratorRunner(adapter, event_store)
await runner.execute_seed(seed, parallel=True) # each AC runs in its own Claude session
Summary
- DependencyAnalyzer in
src/ouroboros/orchestrator/dependency_analyzer.pygenerates execution levels where ACs within each level run in parallel. - PALRouter in
src/ouroboros/routing/router.pyevaluates task complexity and assigns each AC to a specific model tier (Frugal, Standard, or Frontier). - ParallelACExecutor spawns isolated async tasks for each AC, creating separate provider sessions that allow different models to run simultaneously.
- LevelCoordinator resolves file conflicts post-execution without interrupting the parallel processing of other ACs.
- Model configurations in
src/ouroboros/config/models.pymap abstract tiers to concrete endpoints likegpt-4o-mini,claude-sonnet-4-6, andgemini-2.5-pro.
Frequently Asked Questions
How does Ouroboros decide which LLM model to use for each Acceptance Criterion?
The PALRouter analyzes task complexity based on token count, tool dependencies, and AC depth, then assigns a tier (Frugal, Standard, or Frontier). This tier resolves to a specific model defined in src/ouroboros/config/models.py, such as routing lightweight tasks to gpt-4o-mini and complex reasoning tasks to gemini-2.5-pro.
Can different ACs in the same parallel batch use different model providers?
Yes. Because the ClaudeAgentAdapter receives the model name dynamically from the router and creates a separate session for each AC, a single parallel batch can simultaneously execute ACs through OpenAI, Anthropic, and Google endpoints. Each maintains its own isolated connection without cross-provider interference.
What happens if two parallel ACs modify the same file?
The LevelCoordinator detects file-modifying tool call conflicts after a level completes. It launches a dedicated Coordinator Claude session to reconcile the changes before proceeding to the next dependency level. This resolution occurs independently and does not block or restart the other parallel AC executions.
How does the PAL router determine the appropriate model tier?
The router evaluates three specific metrics: the token count of the task context, the tool dependencies required (such as file system or database access), and the AC depth within the task decomposition hierarchy. These factors collectively determine whether an AC receives Frugal, Standard, or Frontier tier allocation according to the logic in src/ouroboros/routing/router.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →