# Ouroboros 3-Stage Evaluation Pipeline: Mechanical → Semantic → Multi-Model Consensus Explained

> Discover Ouroboros's 3-stage evaluation pipeline: Mechanical, Semantic, and Multi-Model Consensus. Filter defective code efficiently before costly LLM judges.

- Repository: [Q00/ouroboros](https://github.com/Q00/ouroboros)
- Tags: deep-dive
- Published: 2026-03-14

---

**Ouroboros evaluates every generated artifact through a progressive three-stage pipeline—Mechanical Verification, Semantic Evaluation, and Multi-Model Consensus—that filters out defective code at zero cost before escalating to frontier-tier LLM judges only when necessary.**

The Q00/ouroboros repository implements a rigorous evaluation architecture designed to balance computational cost with high-quality code assessment. This 3-stage evaluation process progressively applies mechanical checks, semantic analysis against acceptance criteria, and optional multi-model consensus voting to ensure only valid, goal-aligned artifacts reach final approval.

## Stage 1: Mechanical Verification

The first stage acts as a cost-saving gatekeeper, executing stateless, zero-cost checks before any expensive LLM inference occurs.

### Implementation in [`mechanical.py`](https://github.com/Q00/ouroboros/blob/main/mechanical.py)

Located in [`src/ouroboros/evaluation/mechanical.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/mechanical.py), the `MechanicalVerifier` class receives a `MechanicalConfig` and runs a configurable list of `CheckType` validations. The core verification loop executes shell commands for linting, building, testing, static analysis, and optional coverage checks via the `run_command` utility.

```python

# Conceptual usage of MechanicalVerifier

from ouroboros.evaluation.mechanical import MechanicalVerifier, MechanicalConfig

config = MechanicalConfig(
    checks=["lint", "build", "test", "static_analysis"],
    coverage_threshold=80.0
)
verifier = MechanicalVerifier(config)
result = await verifier.verify(artifact_path)

```

The verifier returns a `MechanicalResult` containing a boolean `passed` field, per-check `CheckResult` details, and an optional coverage score. Throughout execution, it emits structured events—`evaluation.stage1.started` and `evaluation.stage1.completed`—through the `create_stage1_started_event` and `create_stage1_completed_event` factories defined in [`src/ouroboros/events/evaluation.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/events/evaluation.py).

### Zero-Cost Guarantee

Because Mechanical Verification relies entirely on local shell execution and static analysis tools, this stage operates at **$0** cost while guaranteeing that code actually compiles, passes tests, and meets coverage thresholds before proceeding to semantic evaluation.

## Stage 2: Semantic Evaluation

Once mechanical checks pass, the pipeline escalates to LLM-based semantic analysis to assess whether the artifact meets its acceptance criteria and aligns with the stated goal.

### LLM Adapter and Prompt Construction

Implemented in [`src/ouroboros/evaluation/semantic.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/semantic.py), the `SemanticEvaluator` utilizes a `LLMAdapter` (backed by LiteLLM) and a `SemanticConfig` that defaults to the `claude-opus-4-6` model. The `build_evaluation_prompt` function constructs a detailed prompt injecting the acceptance criterion, goal, constraints, artifact content, and any source files into the context.

```python
from ouroboros.evaluation.semantic import SemanticEvaluator, SemanticConfig

config = SemanticConfig(
    model="claude-opus-4-6",
    temperature=0.2,
    satisfaction_threshold=0.8
)
evaluator = SemanticEvaluator(llm_adapter, config)
semantic_result = await evaluator.evaluate(context)

```

### Structured JSON Response Handling

The LLM is forced to return a JSON object matching the `SEMANTIC_RESULT_SCHEMA` defined in the source. The `parse_semantic_response` helper validates and normalizes fields into a `SemanticResult` dataclass, measuring **acceptance-criterion compliance**, **goal alignment**, and **drift** (deviation from original intent). The default satisfaction threshold is **0.8**; failure below this threshold aborts the pipeline before Stage 3.

## Stage 3: Multi-Model Consensus

The final stage serves as a frontier-tier safety net, employing multiple independent models to resolve high-stakes or ambiguous evaluations.

### Simple Majority Voting

Defined in [`src/ouroboros/evaluation/consensus.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/consensus.py), the `ConsensusEvaluator` queries three Frontier-tier models—defaulting to GPT-4o, Claude-Opus-4-6, and Gemini-2.5-pro. Each model returns a JSON vote containing `approved`, `confidence`, and `reasoning` fields, validated by `parse_vote_response`. The evaluator calculates a majority ratio against a configurable threshold (default **2/3**), returning a `ConsensusResult` containing the vote list, final ratio, and any dissenting reasoning.

```python
from ouroboros.evaluation.consensus import ConsensusEvaluator, ConsensusConfig

config = ConsensusConfig(
    models=["gpt-4o", "claude-opus-4-6", "gemini-2.5-pro"],
    majority_threshold=0.67
)
consensus = ConsensusEvaluator(llm_adapter, config)
result = await consensus.evaluate(context, semantic_result)

```

### Deliberative Consensus Mode

For high-risk decisions, Ouroboros provides `DeliberativeConsensus`, which extends simple voting with an **Advocate**, **Devil's Advocate** (performing ontological analysis), and a **Judge** role. This deliberative variant ensures the solution addresses the root cause rather than surface-level symptoms, adding ontological reasoning to the approval process.

## Pipeline Orchestration and Trigger Logic

The `EvaluationPipeline` class in [`src/ouroboros/evaluation/pipeline.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/pipeline.py) orchestrates the three stages through a strict sequential flow:

1. **Execute Mechanical** via `self._mechanical.verify`. If checks fail, abort immediately.
2. **Execute Semantic** via `self._semantic.evaluate`. If the acceptance criterion is not met, abort.
3. **Evaluate Trigger** by building a `TriggerContext` from the semantic result and querying `self._trigger.evaluate` (implemented in [`src/ouroboros/evaluation/trigger.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/trigger.py)).
4. **Conditional Consensus**: Only if the trigger matrix returns `should_trigger = True` does the pipeline invoke `self._consensus.evaluate`.

The final `EvaluationResult` aggregates all stage results, identifies the highest completed stage, and sets a `final_approved` flag based on the cumulative outcome.

```python
from ouroboros.evaluation.pipeline import run_evaluation_pipeline, PipelineConfig
from ouroboros.evaluation.models import EvaluationContext
from ouroboros.providers.litellm_adapter import LiteLLMAdapter

# Build execution context

context = EvaluationContext(
    execution_id="run-001",
    artifact=Path("solution.py").read_text(),
    artifact_type="python",
    current_ac="Implement a thread-safe LRU cache with O(1) operations.",
    goal="Provide production-ready, well-tested code.",
    constraints=["Python 3.10+", "No external dependencies"],
)

# Run full pipeline

llm_adapter = LiteLLMAdapter()
result = await run_evaluation_pipeline(context, llm_adapter)

# Inspect results

print(f"Final approved: {result.value.final_approved}")
print(f"Highest stage: {result.value.highest_stage}")

```

## Summary

- **Mechanical Verification** ([`src/ouroboros/evaluation/mechanical.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/mechanical.py)) provides zero-cost validation through linting, testing, and coverage checks, aborting early on failure.
- **Semantic Evaluation** ([`src/ouroboros/evaluation/semantic.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/semantic.py)) uses a Standard-tier LLM (default `claude-opus-4-6`) to score acceptance-criterion compliance and goal alignment with a default threshold of 0.8.
- **Multi-Model Consensus** ([`src/ouroboros/evaluation/consensus.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/consensus.py)) optionally invokes Frontier-tier models (GPT-4o, Claude-Opus-4-6, Gemini-2.5-pro) for majority voting or deliberative review, triggered only when the semantic result meets specific risk criteria.
- **Orchestration** ([`src/ouroboros/evaluation/pipeline.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/pipeline.py)) handles stage sequencing, early termination, and event emission through `run_evaluation_pipeline`.

## Frequently Asked Questions

### What happens if Mechanical Verification fails?

If the `MechanicalVerifier` detects lint errors, build failures, or insufficient test coverage, it returns a `MechanicalResult` with `passed = False`. The `EvaluationPipeline` aborts immediately, returning a final result with `final_approved = False` and `highest_stage = 1`, preventing wasted LLM inference costs on broken code.

### How does the Consensus Trigger decide when to run Stage 3?

The `ConsensusTrigger` in [`src/ouroboros/evaluation/trigger.py`](https://github.com/Q00/ouroboros/blob/main/src/ouroboros/evaluation/trigger.py) evaluates a `TriggerContext` containing semantic scores, confidence levels, and artifact metadata. It returns `should_trigger = True` when semantic satisfaction scores fall within a configurable uncertainty band (e.g., 0.7–0.85) or when the artifact type requires extra scrutiny, ensuring expensive frontier-tier votes only occur for ambiguous cases.

### What is the difference between Simple Voting and Deliberative Consensus?

**Simple Voting** (`ConsensusEvaluator`) aggregates independent binary approvals from three models and requires a 2/3 majority. **Deliberative Consensus** (`DeliberativeConsensus`) adds structured debate roles—an Advocate promoting the solution, a Devil's Advocate challenging underlying assumptions with ontological analysis, and a Judge synthesizing the arguments—to ensure root-cause alignment before approval.

### How much does the evaluation pipeline cost to run?

Stage 1 (Mechanical) costs **$0** as it uses only local tooling. Stage 2 (Semantic) incurs a single Standard-tier LLM call (~$0.03–$0.15 depending on context length). Stage 3 (Consensus) only runs when triggered, costing 3× Frontier-tier tokens (~$0.30–$1.50) for simple voting or more for deliberative rounds. Early aborts at Stage 1 or 2 prevent unnecessary Stage 3 expenses.