Ouroboros Ambiguity Score Calculation and 0.2 Threshold Gate Explained
The Ouroboros Ambiguity Score is computed as 1 minus the weighted average of clarity component scores (goal, constraints, success criteria, and optional codebase context), rounded to four decimal places, and the 0.2 threshold gate enforced by AMBIGUITY_THRESHOLD blocks Seed generation unless the overall score is ≤ 0.2.
The Q00/ouroboros repository implements a deterministic requirements clarification pipeline that quantifies specification uncertainty through the Ambiguity Score calculation before permitting AI-generated code seeds. This metric ensures that only requirements meeting a strict clarity standard—governed by the 0.2 threshold gate—progress to the Seed generation phase, preventing wasted computation on vague or incomplete specifications.
How the Ambiguity Score Calculation Works
The Mathematical Formula
The calculation inverts clarity into ambiguity. First, the system collects clarity scores (0.0 = totally unclear, 1.0 = perfectly clear) from an LLM for each requirement component. Then, AmbiguityScorer._calculate_overall_score in src/ouroboros/bigbang/ambiguity.py (lines 474‑490) applies the formula:
overall_ambiguity = 1.0 - (goal_clarity × weight_goal + constraint_clarity × weight_constraint + success_clarity × weight_success + [context_clarity × weight_context])
The result is rounded to four decimal places. A score of 0.0 indicates perfectly clear requirements, while 1.0 indicates total ambiguity.
Component Weights and Project Types
Weights differ based on whether the project is greenfield (new codebase) or brownfield (existing codebase), as defined in src/ouroboros/bigbang/ambiguity.py (lines 31‑40):
Greenfield weights:
- Goal: 0.40
- Constraint: 0.30
- Success Criteria: 0.30
Brownfield weights:
- Goal: 0.35
- Constraint: 0.25
- Success Criteria: 0.25
- Context (existing code): 0.15
Brownfield projects include the additional Context component to account for codebase clarity, drawn from analysis in src/ouroboros/bigbang/interview.py.
The 0.2 Threshold Gate Mechanism
Gate Implementation in ambiguity.py
The threshold is hardcoded as the constant AMBIGUITY_THRESHOLD = 0.2 in src/ouroboros/bigbang/ambiguity.py (lines 28‑31). This value represents Non‑Functional Requirement 6 (NFR6), mandating that only requirements with ≤ 20% ambiguity may proceed to Seed generation.
The gate logic resides in the AmbiguityScore dataclass via the is_ready_for_seed property (lines 108‑115):
@dataclass(frozen=True, slots=True)
class AmbiguityScore:
overall_score: float
breakdown: ScoreBreakdown
@property
def is_ready_for_seed(self) -> bool:
return self.overall_score <= AMBIGUITY_THRESHOLD
A standalone helper function is_ready_for_seed(score: AmbiguityScore) -> bool provides the same check for functional programming contexts throughout the engine.
Downstream Pipeline Impact
When is_ready_for_seed returns False, the orchestration layer in src/ouroboros/evaluation/pipeline.py triggers the clarification loop instead of Seed generation. The generate_clarification_questions method (lines 492‑527 in ambiguity.py) analyzes low‑scoring components to generate targeted follow‑up questions. Only when the score drops to ≤ 0.2 does the workflow advance to SeedGenerator.
Code Implementation Details
Core Data Structures
The scoring pipeline relies on immutable data structures defined in src/ouroboros/bigbang/ambiguity.py:
@dataclass(frozen=True)
class ScoreBreakdown:
goal_clarity_score: float
goal_clarity_justification: str
constraint_clarity_score: float
constraint_clarity_justification: str
success_criteria_clarity_score: float
success_criteria_justification: str
# Brownfield only:
context_clarity_score: Optional[float] = None
context_clarity_justification: Optional[str] = None
Scoring Workflow Steps
- Context Building:
_build_interview_context(lines 404‑420) flattens theInterviewStateinto a transcript string. - LLM Prompting:
_build_scoring_system_promptand_build_scoring_user_prompt(lines 322‑380) request JSON‑formatted clarity ratings. - Parsing:
_parse_scoring_response(lines 378‑460) clamps values to[0.0, 1.0]and constructs aScoreBreakdown. - Calculation:
_calculate_overall_scoreapplies the weighted formula. - Gating:
is_ready_for_seedcompares againstAMBIGUITY_THRESHOLD.
Practical Example: Computing the Score
Below is a runnable example demonstrating the calculation and gate logic:
from ouroboros.bigbang.ambiguity import (
AmbiguityScorer,
AmbiguityScore,
is_ready_for_seed,
AMBIGUITY_THRESHOLD
)
from ouroboros.bigbang.interview import InterviewState, RoundData
# Construct a minimal interview state
state = InterviewState(
interview_id="demo-123",
initial_context="Build a REST API",
rounds=[
RoundData(question="What framework?", user_response="FastAPI"),
RoundData(question="Constraints?", user_response="Async only"),
],
is_brownfield=False, # Uses greenfield weights: 0.4, 0.3, 0.3
)
# Simulate parsed LLM response (clarity scores)
scorer = AmbiguityScorer.__new__(AmbiguityScorer) # bypass init for demo
breakdown = scorer._parse_scoring_response('''
{
"goal_clarity_score": 0.95,
"goal_clarity_justification": "Specific API goal stated",
"constraint_clarity_score": 0.80,
"constraint_clarity_justification": "Async constraint clear",
"success_criteria_clarity_score": 0.85,
"success_criteria_justification": "Testing criteria defined"
}
''')
# Calculate overall ambiguity
overall = scorer._calculate_overall_score(breakdown)
score = AmbiguityScore(overall_score=overall, breakdown=breakdown)
print(f"Overall Ambiguity Score: {score.overall_score:.4f}")
print(f"Threshold: {AMBIGUITY_THRESHOLD}")
print(f"Ready for Seed: {score.is_ready_for_seed}") # True if ≤ 0.2
With the clarity scores above (0.95, 0.80, 0.85), the weighted clarity is (0.95×0.4)+(0.80×0.3)+(0.85×0.3)=0.875, yielding an ambiguity score of 0.1250. Since 0.1250 ≤ 0.2, is_ready_for_seed returns True.
Summary
- The Ambiguity Score calculation in Q00/ouroboros derives from a weighted average of component clarity scores, inverted to represent uncertainty (1 – weighted_clarity).
- Greenfield projects use weights of 40/30/30 for goal/constraints/success criteria, while brownfield adds a 15% weight for codebase context.
- The 0.2 threshold gate is implemented via the constant
AMBIGUITY_THRESHOLDinsrc/ouroboros/bigbang/ambiguity.pyand enforced by theis_ready_for_seedproperty. - Scores ≤ 0.2 unlock Seed generation; higher scores trigger the clarification question loop via
generate_clarification_questions. - The pipeline relies on strict JSON parsing, value clamping, and four‑decimal rounding to ensure deterministic, reproducible gate decisions.
Frequently Asked Questions
What components are evaluated in the Ouroboros Ambiguity Score calculation?
The calculation evaluates goal clarity, constraint clarity, and success criteria clarity for all projects. Brownfield projects (existing codebases) include a fourth component: codebase context clarity. Each component is rated 0.0 (unclear) to 1.0 (clear) by an LLM before being weighted and combined into the final score.
How does the 0.2 threshold gate function in the pipeline?
The gate functions as a hard boundary defined by AMBIGUITY_THRESHOLD = 0.2 in src/ouroboros/bigbang/ambiguity.py. The AmbiguityScore.is_ready_for_seed property returns True only when overall_score <= 0.2. If the gate returns False, the engine enters a clarification loop that generates additional questions based on low‑scoring components, preventing premature Seed generation.
Where is the Ambiguity Score calculation implemented in the source code?
The primary implementation resides in src/ouroboros/bigbang/ambiguity.py. Key methods include AmbiguityScorer._calculate_overall_score (lines 474‑490) for the mathematical calculation, _parse_scoring_response (lines 378‑460) for processing LLM output, and the is_ready_for_seed property (lines 108‑115) for threshold enforcement. The orchestration logic is located in src/ouroboros/evaluation/pipeline.py.
Can the ambiguity threshold be adjusted for different strictness levels?
Yes. The threshold is controlled by the constant AMBIGUITY_THRESHOLD defined at line 28 in src/ouroboros/bigbang/ambiguity.py. Modifying this single value tightens or loosens the gate across the entire platform without requiring changes to downstream logic. However, the default value of 0.2 is tied to NFR6 and should be adjusted only after validating the impact on Seed quality and clarification loop frequency.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →