How the Constraint Checking System Works in Local Deep Research: DualConfidenceChecker, ThresholdChecker, and StrictChecker Explained

The constraint checking system in learningcircuit/local-deep-research uses an extensible hierarchy where DualConfidenceChecker analyzes positive, negative, and uncertainty signals with re-evaluation loops, ThresholdChecker applies simple yes/no satisfaction scoring against global rates, and StrictChecker enforces a high-confidence bar (default 0.9) that rejects candidates on any single failure.

The local-deep-research repository implements a modular constraint checking system that evaluates search candidates against user-defined requirements using three distinct strategies. Each checker inherits from a common base class but implements different scoring algorithms to balance speed, nuance, and strictness.

The Base Architecture

All constraint checkers inherit from BaseConstraintChecker in src/local_deep_research/advanced_search_system/constraint_checking/base_constraint_checker.py. This abstract base defines the shared interface, manages the LLM model (BaseChatModel), handles evidence gathering, and calculates weighted scores.

The common workflow follows four steps:

  1. Instantiation – Subclasses receive a LangChain BaseChatModel instance and an evidence_gatherer callable provided by the search strategy.
  2. Evaluation – Calling check_candidate(candidate, constraints) triggers the concrete scoring logic.
  3. Evidence Processing – Each constraint is processed via _gather_evidence_for_constraint, which retrieves candidate-specific evidence. The checker then produces a score and evaluates should_reject_candidate.
  4. Result Packaging – Scores are stored in a ConstraintCheckResult dataclass containing the candidate, total score, per-constraint breakdown, rejection flag, reason, and detailed results.

DualConfidenceChecker: Nuanced Confidence Analysis

Located in dual_confidence_checker.py, the DualConfidenceChecker performs sophisticated three-way confidence analysis for each constraint, distinguishing between supporting, contradicting, and uncertain evidence.

Scoring Methodology

For every constraint, the checker calls EvidenceAnalyzer.analyze_evidence_dual_confidence to extract positive, negative, and uncertainty components. The final score is derived from these weighted components rather than a single yes/no response.

Re-evaluation Loops

If the overall uncertainty exceeds uncertainty_threshold (default 0.6), the checker automatically triggers _evaluate_constraint_with_reevaluation. This repeats evidence gathering up to max_reevaluations times to resolve ambiguous cases before making a final decision.

Configurable Thresholds

Rejection decisions use three configurable thresholds managed by should_reject_candidate_from_averages:

  • negative_threshold (default 0.25): Reject when negative evidence exceeds 25% of the total
  • positive_threshold (default 0.4): Reject when positive evidence falls below 40%
  • uncertainty_penalty: Dynamically lowers the final score when uncertainty remains high after re-evaluation

The total score is calculated via _calculate_weighted_score, which aggregates dual-confidence scores across all constraints.

ThresholdChecker: Simple Satisfaction Scoring

The ThresholdChecker in threshold_checker.py provides a faster, less nuanced alternative that requests a single satisfaction score per constraint.

Binary Evaluation

The method _check_constraint_satisfaction builds a concise prompt requesting a numeric satisfaction score (0-1) from the LLM. The response is parsed directly without complex multi-component analysis.

Global Rejection Rules

This checker uses two thresholds for final decisions:

  • satisfaction_threshold (default 0.7): Minimum score for a single constraint to be considered satisfied
  • required_satisfaction_rate (default 0.8): Proportion of constraints that must be satisfied overall

If satisfaction_rate < self.required_satisfaction_rate, the candidate is rejected immediately. The total score equals the satisfaction rate (or 0 on rejection).

StrictChecker: High-Confidence Enforcement

The StrictChecker in strict_checker.py implements a hard-stop evaluation requiring every constraint to meet a very high confidence bar.

Strict Scoring Logic

The _evaluate_constraint_strictly method applies different logic based on constraint type:

  • For NAME_PATTERN constraints, _check_name_pattern_strictly applies deterministic pattern checks and requires a ≥ 0.95 score when name_pattern_required is enabled.
  • For standard constraints, a strict prompt asks the LLM to return high scores only when evidence is "clear, unambiguous," with numeric extraction from the reply.

Zero-Tolerance Thresholds

  • strict_threshold (default 0.9): Global high bar every constraint must cross
  • Immediate rejection occurs if any constraint fails the threshold or if evidence is missing

Unlike the other checkers, StrictChecker does not use weighted averaging. The total score is binary: 1 if all constraints pass, 0 if any fail.

Integration with Search Strategies

All three checkers plug into search strategies via the evidence_gatherer argument. Strategies such as DualConfidenceStrategy, ThresholdStrategy, and StrictStrategy determine what evidence to collect (search engines, local embeddings, etc.), then delegate evaluation to the appropriate checker.


# Example: Using DualConfidenceChecker in a strategy

from langchain_openai import OpenAI
from local_deep_research.advanced_search_system.constraint_checking.dual_confidence_checker import DualConfidenceChecker
from local_deep_research.advanced_search_system.strategies.dual_confidence_strategy import DualConfidenceStrategy

# 1️⃣ Initialise the LLM and evidence gatherer (provided by the strategy)

model = OpenAI(api_key="…")                         # <- placeholder

strategy = DualConfidenceStrategy(model=model)

# 2️⃣ Build the checker (inherits the model & gatherer from the strategy)

checker = DualConfidenceChecker(
    model=model,
    evidence_gatherer=strategy.gather_evidence,
    negative_threshold=0.2,
    positive_threshold=0.5,
    uncertainty_threshold=0.55,
    max_reevaluations=3,
)

# 3️⃣ Run the check

result = checker.check_candidate(candidate, constraints)

print("Reject?" , result.should_reject)
print("Total score:", result.total_score)
print("Details:", result.detailed_results)

# Example: ThresholdChecker (simpler, faster)

from local_deep_research.advanced_search_system.constraint_checking.threshold_checker import ThresholdChecker
checker = ThresholdChecker(
    model=model,
    evidence_gatherer=strategy.gather_evidence,
    satisfaction_threshold=0.75,
    required_satisfaction_rate=0.85,
)

result = checker.check_candidate(candidate, constraints)

# Example: StrictChecker (hard‑stop on any low confidence)

from local_deep_research.advanced_search_system.constraint_checking.strict_checker import StrictChecker
checker = StrictChecker(
    model=model,
    evidence_gatherer=strategy.gather_evidence,
    strict_threshold=0.92,
    name_pattern_required=True,
)

result = checker.check_candidate(candidate, constraints)

All three checkers return a ConstraintCheckResult object containing the candidate, total score, per-constraint scores, rejection status, reason, and detailed results suitable for logging or UI display.

Summary

  • BaseConstraintChecker provides the abstract foundation managing LLM instances, evidence gathering, and result packaging in base_constraint_checker.py.
  • DualConfidenceChecker analyzes positive, negative, and uncertainty signals with configurable thresholds (0.25 negative, 0.4 positive) and automatic re-evaluation loops when uncertainty exceeds 0.6.
  • ThresholdChecker uses simple 0-1 satisfaction scores per constraint and rejects candidates failing to meet an 80% global satisfaction rate by default.
  • StrictChecker enforces a 0.9 minimum score on every constraint with special handling for NAME_PATTERN types, producing binary pass/fail results.
  • All checkers integrate with search strategies through the evidence_gatherer callable, returning standardized ConstraintCheckResult objects.

Frequently Asked Questions

What is the difference between DualConfidenceChecker and ThresholdChecker?

DualConfidenceChecker breaks evidence into positive, negative, and uncertainty components, allowing nuanced scoring and re-evaluation when ambiguity is high. ThresholdChecker requests a single satisfaction score per constraint and evaluates against global satisfaction rates, making it faster but less granular regarding conflicting evidence.

How does StrictChecker handle name pattern constraints differently?

According to the source code in strict_checker.py, when name_pattern_required is enabled, StrictChecker calls _check_name_pattern_strictly to apply deterministic validation and requires a score of at least 0.95 for these specific constraints, higher than the standard 0.9 strict_threshold used for other constraint types.

Can I configure the thresholds for rejection in each checker?

Yes. DualConfidenceChecker accepts negative_threshold, positive_threshold, and uncertainty_threshold parameters. ThresholdChecker exposes satisfaction_threshold and required_satisfaction_rate. StrictChecker allows adjustment of strict_threshold and name_pattern_required flags to customize strictness levels.

What happens when DualConfidenceChecker encounters uncertain evidence?

When uncertainty exceeds the default 0.6 threshold, DualConfidenceChecker triggers _evaluate_constraint_with_reevaluation to gather additional evidence up to max_reevaluations times. If uncertainty remains high, the uncertainty_penalty reduces the final score, potentially triggering rejection if scores fall below the configured positive or negative thresholds.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →