Ouroboros 3-Stage Evaluation Pipeline: Mechanical → Semantic → Multi-Model Consensus Explained

Ouroboros evaluates every generated artifact through a progressive three-stage pipeline—Mechanical Verification, Semantic Evaluation, and Multi-Model Consensus—that filters out defective code at zero cost before escalating to frontier-tier LLM judges only when necessary.

The Q00/ouroboros repository implements a rigorous evaluation architecture designed to balance computational cost with high-quality code assessment. This 3-stage evaluation process progressively applies mechanical checks, semantic analysis against acceptance criteria, and optional multi-model consensus voting to ensure only valid, goal-aligned artifacts reach final approval.

Stage 1: Mechanical Verification

The first stage acts as a cost-saving gatekeeper, executing stateless, zero-cost checks before any expensive LLM inference occurs.

Implementation in mechanical.py

Located in src/ouroboros/evaluation/mechanical.py, the MechanicalVerifier class receives a MechanicalConfig and runs a configurable list of CheckType validations. The core verification loop executes shell commands for linting, building, testing, static analysis, and optional coverage checks via the run_command utility.


# Conceptual usage of MechanicalVerifier

from ouroboros.evaluation.mechanical import MechanicalVerifier, MechanicalConfig

config = MechanicalConfig(
    checks=["lint", "build", "test", "static_analysis"],
    coverage_threshold=80.0
)
verifier = MechanicalVerifier(config)
result = await verifier.verify(artifact_path)

The verifier returns a MechanicalResult containing a boolean passed field, per-check CheckResult details, and an optional coverage score. Throughout execution, it emits structured events—evaluation.stage1.started and evaluation.stage1.completed—through the create_stage1_started_event and create_stage1_completed_event factories defined in src/ouroboros/events/evaluation.py.

Zero-Cost Guarantee

Because Mechanical Verification relies entirely on local shell execution and static analysis tools, this stage operates at $0 cost while guaranteeing that code actually compiles, passes tests, and meets coverage thresholds before proceeding to semantic evaluation.

Stage 2: Semantic Evaluation

Once mechanical checks pass, the pipeline escalates to LLM-based semantic analysis to assess whether the artifact meets its acceptance criteria and aligns with the stated goal.

LLM Adapter and Prompt Construction

Implemented in src/ouroboros/evaluation/semantic.py, the SemanticEvaluator utilizes a LLMAdapter (backed by LiteLLM) and a SemanticConfig that defaults to the claude-opus-4-6 model. The build_evaluation_prompt function constructs a detailed prompt injecting the acceptance criterion, goal, constraints, artifact content, and any source files into the context.

from ouroboros.evaluation.semantic import SemanticEvaluator, SemanticConfig

config = SemanticConfig(
    model="claude-opus-4-6",
    temperature=0.2,
    satisfaction_threshold=0.8
)
evaluator = SemanticEvaluator(llm_adapter, config)
semantic_result = await evaluator.evaluate(context)

Structured JSON Response Handling

The LLM is forced to return a JSON object matching the SEMANTIC_RESULT_SCHEMA defined in the source. The parse_semantic_response helper validates and normalizes fields into a SemanticResult dataclass, measuring acceptance-criterion compliance, goal alignment, and drift (deviation from original intent). The default satisfaction threshold is 0.8; failure below this threshold aborts the pipeline before Stage 3.

Stage 3: Multi-Model Consensus

The final stage serves as a frontier-tier safety net, employing multiple independent models to resolve high-stakes or ambiguous evaluations.

Simple Majority Voting

Defined in src/ouroboros/evaluation/consensus.py, the ConsensusEvaluator queries three Frontier-tier models—defaulting to GPT-4o, Claude-Opus-4-6, and Gemini-2.5-pro. Each model returns a JSON vote containing approved, confidence, and reasoning fields, validated by parse_vote_response. The evaluator calculates a majority ratio against a configurable threshold (default 2/3), returning a ConsensusResult containing the vote list, final ratio, and any dissenting reasoning.

from ouroboros.evaluation.consensus import ConsensusEvaluator, ConsensusConfig

config = ConsensusConfig(
    models=["gpt-4o", "claude-opus-4-6", "gemini-2.5-pro"],
    majority_threshold=0.67
)
consensus = ConsensusEvaluator(llm_adapter, config)
result = await consensus.evaluate(context, semantic_result)

Deliberative Consensus Mode

For high-risk decisions, Ouroboros provides DeliberativeConsensus, which extends simple voting with an Advocate, Devil's Advocate (performing ontological analysis), and a Judge role. This deliberative variant ensures the solution addresses the root cause rather than surface-level symptoms, adding ontological reasoning to the approval process.

Pipeline Orchestration and Trigger Logic

The EvaluationPipeline class in src/ouroboros/evaluation/pipeline.py orchestrates the three stages through a strict sequential flow:

  1. Execute Mechanical via self._mechanical.verify. If checks fail, abort immediately.
  2. Execute Semantic via self._semantic.evaluate. If the acceptance criterion is not met, abort.
  3. Evaluate Trigger by building a TriggerContext from the semantic result and querying self._trigger.evaluate (implemented in src/ouroboros/evaluation/trigger.py).
  4. Conditional Consensus: Only if the trigger matrix returns should_trigger = True does the pipeline invoke self._consensus.evaluate.

The final EvaluationResult aggregates all stage results, identifies the highest completed stage, and sets a final_approved flag based on the cumulative outcome.

from ouroboros.evaluation.pipeline import run_evaluation_pipeline, PipelineConfig
from ouroboros.evaluation.models import EvaluationContext
from ouroboros.providers.litellm_adapter import LiteLLMAdapter

# Build execution context

context = EvaluationContext(
    execution_id="run-001",
    artifact=Path("solution.py").read_text(),
    artifact_type="python",
    current_ac="Implement a thread-safe LRU cache with O(1) operations.",
    goal="Provide production-ready, well-tested code.",
    constraints=["Python 3.10+", "No external dependencies"],
)

# Run full pipeline

llm_adapter = LiteLLMAdapter()
result = await run_evaluation_pipeline(context, llm_adapter)

# Inspect results

print(f"Final approved: {result.value.final_approved}")
print(f"Highest stage: {result.value.highest_stage}")

Summary

  • Mechanical Verification (src/ouroboros/evaluation/mechanical.py) provides zero-cost validation through linting, testing, and coverage checks, aborting early on failure.
  • Semantic Evaluation (src/ouroboros/evaluation/semantic.py) uses a Standard-tier LLM (default claude-opus-4-6) to score acceptance-criterion compliance and goal alignment with a default threshold of 0.8.
  • Multi-Model Consensus (src/ouroboros/evaluation/consensus.py) optionally invokes Frontier-tier models (GPT-4o, Claude-Opus-4-6, Gemini-2.5-pro) for majority voting or deliberative review, triggered only when the semantic result meets specific risk criteria.
  • Orchestration (src/ouroboros/evaluation/pipeline.py) handles stage sequencing, early termination, and event emission through run_evaluation_pipeline.

Frequently Asked Questions

What happens if Mechanical Verification fails?

If the MechanicalVerifier detects lint errors, build failures, or insufficient test coverage, it returns a MechanicalResult with passed = False. The EvaluationPipeline aborts immediately, returning a final result with final_approved = False and highest_stage = 1, preventing wasted LLM inference costs on broken code.

How does the Consensus Trigger decide when to run Stage 3?

The ConsensusTrigger in src/ouroboros/evaluation/trigger.py evaluates a TriggerContext containing semantic scores, confidence levels, and artifact metadata. It returns should_trigger = True when semantic satisfaction scores fall within a configurable uncertainty band (e.g., 0.7–0.85) or when the artifact type requires extra scrutiny, ensuring expensive frontier-tier votes only occur for ambiguous cases.

What is the difference between Simple Voting and Deliberative Consensus?

Simple Voting (ConsensusEvaluator) aggregates independent binary approvals from three models and requires a 2/3 majority. Deliberative Consensus (DeliberativeConsensus) adds structured debate roles—an Advocate promoting the solution, a Devil's Advocate challenging underlying assumptions with ontological analysis, and a Judge synthesizing the arguments—to ensure root-cause alignment before approval.

How much does the evaluation pipeline cost to run?

Stage 1 (Mechanical) costs $0 as it uses only local tooling. Stage 2 (Semantic) incurs a single Standard-tier LLM call ($0.03–$0.15 depending on context length). Stage 3 (Consensus) only runs when triggered, costing 3× Frontier-tier tokens ($0.30–$1.50) for simple voting or more for deliberative rounds. Early aborts at Stage 1 or 2 prevent unnecessary Stage 3 expenses.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →