LoopX Testing Strategies: A Multi-Layered Quality System Explained
LoopX testing strategies combine fast deterministic checks, durable public smokes, risk-based canaries, and specialized model-behavior gates to validate code at different distances from shipped behavior.
The LoopX repository implements a sophisticated quality assurance philosophy that prioritizes both developer velocity and production safety. Rather than relying on a single test suite, the system layers multiple validation approaches—each with distinct purposes, cadences, and failure modes. This article examines the architecture of LoopX testing strategies as documented in docs/development/testing-and-quality.md and related source files.
The Seven Layers of LoopX Testing Strategies
LoopX organizes validation into seven complementary layers. A change must survive appropriate layers before reaching production, with deterministic failures always trumping probabilistic passes.
| Layer | Proves | Typical Cadence |
|---|---|---|
| Unit & contract tests | Pure rules, schemas, state transitions, invalid-state rejection | Every PR via python-tests.yml |
| Durable public smokes | CLI & cross-module behavior via public-safe fixtures | Daily or manual full runs |
| Catalog-informed canary | Minimal risk-based slice covering changed surfaces | Pre-merge for sensitive changes |
| CLI output budgets | Agent-facing output stays bounded and visible | PR CI and pre-merge |
| Public-safe decision replay | Quota-to-scheduler path honors invariants | Regression & control-plane changes |
| Model-behavior qualification | Real model interpretation of default packets | Low-frequency shadow gate |
| Release outcome baseline | Stable vs. candidate comparison under matched semantics | Release qualification |
Each layer validates a different "distance" from production. Fast layers catch obvious errors immediately; slower layers catch subtle integration and behavioral issues before customers see them.
Unit and Contract Tests: The Fast PR Gate
The foundation of LoopX testing strategies is deterministic, fast-running validation. According to docs/development/testing-and-quality.md, these tests cover:
- Pure rules—business logic without side effects
- Schemas—data structure validation
- State-transition tables—valid and invalid state progressions
- Invalid-state rejection—error handling paths
These run on every PR via .github/workflows/python-tests.yml. The fast gate includes linting, type checking, and unit tests:
# Fast PR gate (run locally before push)
python -m pip install -e ".[test]"
python -m ruff check tests loopx/canary loopx/control_plane loopx/domain_packs loopx/presentation
python -m mypy
python -m pytest -q # unit & contract tests
git diff --check # lint-free diffs
Contract tests are particularly important—they validate that modules honor their published interfaces, preventing the "dependency drift" that集成 tests often miss.
Durable Public Smokes: Executable Proofs
Smokes in LoopX are thin, executable scripts that prove a public path still honors a durable invariant. The criteria for "good smokes" are strictly defined in docs/development/good-smokes.md.
A valid smoke must protect at least one of:
- Shipped CLI/runtime behavior
- Reusable public contract
- Public/private boundary
- A regression that previously stranded automation
- Representative end-to-end fixture
Smokes use deterministic oracles—independently reviewed rules or protocols—never mere output copying. They also include negative/mutation cases to prove invariants resist rule inversion.
Example focused smoke execution:
# Focused smoke during development
python examples/control_plane/interaction-scheduler-authority-smoke.py
The full public suite runs daily:
# Full-public smoke suite (daily or manual)
loopx canary smoke-suite --suite full-public --jobs 4 --timeout-seconds 120
Catalog-Informed Canary: Risk-Based Selection
The canary system solves a scaling problem: full smoke suites become slow, but skipping tests risks missing broken surfaces. LoopX testing strategies address this through catalog-informed selection.
The canary runner (loopx/canary/ module) maintains a catalog of which smokes exercise which public surfaces. Given a Git diff, it computes the minimal smoke subset covering all changed surfaces:
# Canary selection for a PR
loopx canary premerge --from-git-diff
This keeps CI fast for routine changes while preserving safety for high-risk modifications. High-risk surfaces are explicitly cataloged; the canary prioritizes their coverage.
Specialized Gates: Budgets, Replay, and Models
Beyond core testing, LoopX testing strategies include three specialized validation layers.
CLI Output Budgets
Agent-facing output must stay bounded. Growth is reported and gated, preventing unbounded output expansion that degrades system performance or contract stability.
Public-Safe Decision Replay
For control-plane and regression changes, LoopX can replay the quota-to-scheduler path using an independently reviewed invariant. This validates that resource allocation decisions remain consistent across code versions.
Model-Behavior Qualification
Provider-backed behavior (e.g., Doubao) requires special handling. Real model calls are slow and non-deterministic, so they run in a low-frequency shadow gate:
# Model behavior qualification (low-frequency shadow gate)
python3 scripts/qualify-doubao-model-behavior-live.py \
--qualification-id <public-safe-run-id>
This script (scripts/qualify-doubao-model-behavior-live.py) only executes when an environment key is present, keeping CI fast while still exercising critical model paths.
Release Outcome Baselines
Before any release, LoopX compares stable versus candidate releases under matched semantics. This final validation layer catches subtle behavioral shifts that unit tests and smokes might miss—particularly important for systems with scheduling or resource allocation semantics.
Interaction Principles: How Layers Compose
LoopX testing strategies follow strict interaction rules:
- Deterministic failure > model pass. A failing contract test blocks regardless of smoke results.
- Focused regression > broad sweep. A targeted test pinpointing a broken rule outranks a large smoke suite covering the area indirectly.
- Catalog-driven selection > manual test picking. The canary's automated selection ensures consistent coverage decisions.
These principles prevent the "test pyramid inversion" where slow, flaky tests displace fast, deterministic ones.
Summary
- LoopX testing strategies use seven layered quality gates with distinct purposes and cadences.
- Unit and contract tests provide fast, deterministic PR validation via
python-tests.yml. - Durable public smokes prove CLI and cross-module invariants using deterministic oracles, not output copying.
- The catalog-informed canary selects minimal test subsets based on Git diffs, balancing speed and coverage.
- Specialized gates cover CLI budgets, decision replay, and model behavior qualification.
- Deterministic failures always trump probabilistic passes—the system prioritizes precision over coverage breadth.
Frequently Asked Questions
What makes a smoke "durable" in LoopX?
A durable smoke uses a deterministic oracle—an independently reviewed rule or protocol—as its correctness check, rather than snapshotting current output as expected behavior. According to docs/development/good-smokes.md, durable smokes also include negative and mutation cases to prove invariants resist rule inversion.
How does the canary select which tests to run?
The canary runner (loopx canary premerge --from-git-diff) consults a catalog mapping public surfaces to their covering tests. It computes the minimal subset of smokes that exercise all surfaces changed in the Git diff. This keeps execution fast while ensuring risk coverage for modified code.
Why are model-behavior tests run as a shadow gate?
Real model calls (e.g., to Doubao) are slow, non-deterministic, and require external credentials. Running them in every PR would destroy CI velocity. The shadow gate in scripts/qualify-doubao-model-behavior-live.py executes only when explicitly enabled, providing critical path validation without routine overhead.
Where are LoopX testing strategies documented?
The primary documentation lives in docs/development/testing-and-quality.md. Supplementary guidance for smoke criteria appears in docs/development/good-smokes.md. Implementation details reside in loopx/canary/ for the selection engine and .github/workflows/python-tests.yml for the fast PR gate.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →