Ouroboros 3-Stage Evaluation Pipeline: Mechanical → Semantic → Multi-Model Consensus Explained
Ouroboros evaluates every generated artifact through a progressive three-stage pipeline—Mechanical Verification, Semantic Evaluation, and Multi-Model Consensus—that filters out defective code at zero cost before escalating to frontier-tier LLM judges only when necessary.
The Q00/ouroboros repository implements a rigorous evaluation architecture designed to balance computational cost with high-quality code assessment. This 3-stage evaluation process progressively applies mechanical checks, semantic analysis against acceptance criteria, and optional multi-model consensus voting to ensure only valid, goal-aligned artifacts reach final approval.
Stage 1: Mechanical Verification
The first stage acts as a cost-saving gatekeeper, executing stateless, zero-cost checks before any expensive LLM inference occurs.
Implementation in mechanical.py
Located in src/ouroboros/evaluation/mechanical.py, the MechanicalVerifier class receives a MechanicalConfig and runs a configurable list of CheckType validations. The core verification loop executes shell commands for linting, building, testing, static analysis, and optional coverage checks via the run_command utility.
# Conceptual usage of MechanicalVerifier
from ouroboros.evaluation.mechanical import MechanicalVerifier, MechanicalConfig
config = MechanicalConfig(
checks=["lint", "build", "test", "static_analysis"],
coverage_threshold=80.0
)
verifier = MechanicalVerifier(config)
result = await verifier.verify(artifact_path)
The verifier returns a MechanicalResult containing a boolean passed field, per-check CheckResult details, and an optional coverage score. Throughout execution, it emits structured events—evaluation.stage1.started and evaluation.stage1.completed—through the create_stage1_started_event and create_stage1_completed_event factories defined in src/ouroboros/events/evaluation.py.
Zero-Cost Guarantee
Because Mechanical Verification relies entirely on local shell execution and static analysis tools, this stage operates at $0 cost while guaranteeing that code actually compiles, passes tests, and meets coverage thresholds before proceeding to semantic evaluation.
Stage 2: Semantic Evaluation
Once mechanical checks pass, the pipeline escalates to LLM-based semantic analysis to assess whether the artifact meets its acceptance criteria and aligns with the stated goal.
LLM Adapter and Prompt Construction
Implemented in src/ouroboros/evaluation/semantic.py, the SemanticEvaluator utilizes a LLMAdapter (backed by LiteLLM) and a SemanticConfig that defaults to the claude-opus-4-6 model. The build_evaluation_prompt function constructs a detailed prompt injecting the acceptance criterion, goal, constraints, artifact content, and any source files into the context.
from ouroboros.evaluation.semantic import SemanticEvaluator, SemanticConfig
config = SemanticConfig(
model="claude-opus-4-6",
temperature=0.2,
satisfaction_threshold=0.8
)
evaluator = SemanticEvaluator(llm_adapter, config)
semantic_result = await evaluator.evaluate(context)
Structured JSON Response Handling
The LLM is forced to return a JSON object matching the SEMANTIC_RESULT_SCHEMA defined in the source. The parse_semantic_response helper validates and normalizes fields into a SemanticResult dataclass, measuring acceptance-criterion compliance, goal alignment, and drift (deviation from original intent). The default satisfaction threshold is 0.8; failure below this threshold aborts the pipeline before Stage 3.
Stage 3: Multi-Model Consensus
The final stage serves as a frontier-tier safety net, employing multiple independent models to resolve high-stakes or ambiguous evaluations.
Simple Majority Voting
Defined in src/ouroboros/evaluation/consensus.py, the ConsensusEvaluator queries three Frontier-tier models—defaulting to GPT-4o, Claude-Opus-4-6, and Gemini-2.5-pro. Each model returns a JSON vote containing approved, confidence, and reasoning fields, validated by parse_vote_response. The evaluator calculates a majority ratio against a configurable threshold (default 2/3), returning a ConsensusResult containing the vote list, final ratio, and any dissenting reasoning.
from ouroboros.evaluation.consensus import ConsensusEvaluator, ConsensusConfig
config = ConsensusConfig(
models=["gpt-4o", "claude-opus-4-6", "gemini-2.5-pro"],
majority_threshold=0.67
)
consensus = ConsensusEvaluator(llm_adapter, config)
result = await consensus.evaluate(context, semantic_result)
Deliberative Consensus Mode
For high-risk decisions, Ouroboros provides DeliberativeConsensus, which extends simple voting with an Advocate, Devil's Advocate (performing ontological analysis), and a Judge role. This deliberative variant ensures the solution addresses the root cause rather than surface-level symptoms, adding ontological reasoning to the approval process.
Pipeline Orchestration and Trigger Logic
The EvaluationPipeline class in src/ouroboros/evaluation/pipeline.py orchestrates the three stages through a strict sequential flow:
- Execute Mechanical via
self._mechanical.verify. If checks fail, abort immediately. - Execute Semantic via
self._semantic.evaluate. If the acceptance criterion is not met, abort. - Evaluate Trigger by building a
TriggerContextfrom the semantic result and queryingself._trigger.evaluate(implemented insrc/ouroboros/evaluation/trigger.py). - Conditional Consensus: Only if the trigger matrix returns
should_trigger = Truedoes the pipeline invokeself._consensus.evaluate.
The final EvaluationResult aggregates all stage results, identifies the highest completed stage, and sets a final_approved flag based on the cumulative outcome.
from ouroboros.evaluation.pipeline import run_evaluation_pipeline, PipelineConfig
from ouroboros.evaluation.models import EvaluationContext
from ouroboros.providers.litellm_adapter import LiteLLMAdapter
# Build execution context
context = EvaluationContext(
execution_id="run-001",
artifact=Path("solution.py").read_text(),
artifact_type="python",
current_ac="Implement a thread-safe LRU cache with O(1) operations.",
goal="Provide production-ready, well-tested code.",
constraints=["Python 3.10+", "No external dependencies"],
)
# Run full pipeline
llm_adapter = LiteLLMAdapter()
result = await run_evaluation_pipeline(context, llm_adapter)
# Inspect results
print(f"Final approved: {result.value.final_approved}")
print(f"Highest stage: {result.value.highest_stage}")
Summary
- Mechanical Verification (
src/ouroboros/evaluation/mechanical.py) provides zero-cost validation through linting, testing, and coverage checks, aborting early on failure. - Semantic Evaluation (
src/ouroboros/evaluation/semantic.py) uses a Standard-tier LLM (defaultclaude-opus-4-6) to score acceptance-criterion compliance and goal alignment with a default threshold of 0.8. - Multi-Model Consensus (
src/ouroboros/evaluation/consensus.py) optionally invokes Frontier-tier models (GPT-4o, Claude-Opus-4-6, Gemini-2.5-pro) for majority voting or deliberative review, triggered only when the semantic result meets specific risk criteria. - Orchestration (
src/ouroboros/evaluation/pipeline.py) handles stage sequencing, early termination, and event emission throughrun_evaluation_pipeline.
Frequently Asked Questions
What happens if Mechanical Verification fails?
If the MechanicalVerifier detects lint errors, build failures, or insufficient test coverage, it returns a MechanicalResult with passed = False. The EvaluationPipeline aborts immediately, returning a final result with final_approved = False and highest_stage = 1, preventing wasted LLM inference costs on broken code.
How does the Consensus Trigger decide when to run Stage 3?
The ConsensusTrigger in src/ouroboros/evaluation/trigger.py evaluates a TriggerContext containing semantic scores, confidence levels, and artifact metadata. It returns should_trigger = True when semantic satisfaction scores fall within a configurable uncertainty band (e.g., 0.7–0.85) or when the artifact type requires extra scrutiny, ensuring expensive frontier-tier votes only occur for ambiguous cases.
What is the difference between Simple Voting and Deliberative Consensus?
Simple Voting (ConsensusEvaluator) aggregates independent binary approvals from three models and requires a 2/3 majority. Deliberative Consensus (DeliberativeConsensus) adds structured debate roles—an Advocate promoting the solution, a Devil's Advocate challenging underlying assumptions with ontological analysis, and a Judge synthesizing the arguments—to ensure root-cause alignment before approval.
How much does the evaluation pipeline cost to run?
Stage 1 (Mechanical) costs $0 as it uses only local tooling. Stage 2 (Semantic) incurs a single Standard-tier LLM call ($0.03–$0.15 depending on context length). Stage 3 (Consensus) only runs when triggered, costing 3× Frontier-tier tokens ($0.30–$1.50) for simple voting or more for deliberative rounds. Early aborts at Stage 1 or 2 prevent unnecessary Stage 3 expenses.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →