How LoopX Handles Data Processing: From Ingestion to Kernel State
LoopX processes data through a deterministic control-plane pipeline that ingests raw observations, validates them against typed contracts, transforms them into state transitions, and persists results to a durable kernel with built-in health monitoring and projection capabilities.
LoopX is an open-source control-plane runtime for AI workflows. Understanding how it handles data processing reveals a carefully architected separation between what executes (capabilities), who executes it (providers), and where state lives (the kernel). This article examines the actual source code implementation to show how data flows from raw input to durable, observable output.
Data Processing Architecture in LoopX
The LoopX data processing pipeline follows a canonical flow documented in [docs/architecture.md](https://github.com/huangruiteng/loopx/blob/main/docs/architecture.md):
Agent → Capability → Provider → Kernel → Provider readback → Capability transition → Kernel
This bidirectional flow ensures that every data operation leaves an auditable trace in kernel state while remaining agnostic to the underlying AI provider (Codex App, Claude Code, or custom runners).
Core Components
| Component | Responsibility | Key Source File |
|---|---|---|
| Capability | Defines what processing happens; owns ingestion, validation, and transition logic | [loopx/capabilities/registry.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py) |
| Provider | Executes tooling calls; returns raw outputs for capability consumption | Provider-specific adapters (Codex, Claude, shell) |
| Kernel | Single source of truth for durable state (todos, gates, evidence, quota) |
[loopx/state/model.py](https://github.com/huangruiteng/loopx/blob/main/loopx/state/model.py) |
| Health Module | Emits observable metrics for monitoring and debugging | Capability-specific health modules |
| Projection | Renders kernel state to human-facing surfaces (dashboards, Kanban boards) | [apps/presentation/dashboard/README.md](https://github.com/huangruiteng/loopx/blob/main/apps/presentation/dashboard/README.md) |
The Six-Stage Data Processing Pipeline
Stage 1: Ingestion
Every capability implements its own ingestion logic to parse raw provider output. The Reward Memory capability demonstrates this pattern in [loopx/capabilities/reward_memory/ingestion.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/ingestion.py):
from loopx.capabilities.reward_memory.ingestion import ingest_memory
# Raw payload from external source (e.g., OpenViking corpus)
raw_payload = {
"corpus_id": "proj-123",
"documents": [
{"id": "doc-1", "content": "Implementation notes...", "metadata": {"version": "2.1"}}
]
}
# Parse, enrich with metadata, and prepare for validation
ingest_result = ingest_memory(raw_payload)
The ingestion layer handles:
- Schema-compatible parsing
- Metadata extraction
- Initial error classification
Stage 2: Validation
Capabilities enforce typed contracts and public-private boundaries before any state mutation. The Reward Memory contract ([loopx/capabilities/reward_memory/contract.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/contract.py)) validates that:
- Required fields are present and correctly typed
- Evidence-gate rules are satisfied
- No private data leaks across capability boundaries
if ingest_result.is_valid:
# Proceed to kernel transition
loopx_kernel.apply_transition(ingest_result.transition)
else:
# Validation failures become structured error evidence
loopx_kernel.record_evidence(ingest_result.validation_errors)
Stage 3: Transition
The kernel applies deterministic state transitions through its registry. In [loopx/capabilities/registry.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py), capabilities register their transition handlers:
# Simplified registry pattern
class CapabilityRegistry:
def register(self, capability_id: str, handler: TransitionHandler):
self._handlers[capability_id] = handler
def apply_transition(self, capability_id: str, transition: StateTransition):
handler = self._handlers[capability_id]
updated_state = handler.execute(self._kernel_state, transition)
self._persist(updated_state)
Transitions operate on four core state primitives defined in [loopx/state/model.py](https://github.com/huangruiteng/loopx/blob/main/loopx/state/model.py):
todos: Actionable work itemsgates: Decision checkpoints requiring human or automated approvalevidence: Immutable records of what occurredquota: Resource consumption tracking
Stage 4: Health Monitoring
Every capability emits health packets for observability. The Reward Memory health module ([loopx/capabilities/reward_memory/health.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/health.py)) generates structured diagnostics:
from loopx.capabilities.reward_memory.health import health_report
metrics = health_report(ingest_result)
print(metrics)
# Output:
# {
# 'pipeline': {
# 'corpus_present': True,
# 'index_present': True,
# 'document_count': 147,
# 'last_ingestion_at': '2025-01-15T09:23:11Z'
# }
# }
Health checks enable:
- Real-time pipeline monitoring
- Automated alerting on data quality degradation
- Debugging support for multi-turn AI workflows
Stage 5: Persistence
The kernel serializes state to .loopx/registry.json and optional project-scoped files. This dual-write pattern ensures:
- Global consistency: The registry maintains cross-project references
- Local reproducibility: Project files enable standalone execution
Persistence is atomic: transitions either fully commit or roll back, preventing partial state corruption during provider interruptions.
Stage 6: Projection
Processed data becomes human-readable through projection sinks. The dashboard application ([apps/presentation/dashboard/README.md](https://github.com/huangruiteng/loopx/blob/main/apps/presentation/dashboard/README.md)) renders kernel state as:
- Interactive Kanban boards (Lark integration)
- HTML reports with embedded evidence
- Real-time operational metrics
Projection decouples data processing from presentation, allowing the same kernel state to serve multiple audience needs.
Complete Example: Content-Ops Item Lifecycle
The Content Ops capability demonstrates the full pipeline for markdown content processing in [loopx/capabilities/content_ops/item_lifecycle.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/content_ops/item_lifecycle.py):
from loopx.capabilities.content_ops.item_lifecycle import ItemLifecycle
from loopx.capabilities.content_ops.schemas import ContentItem
# Raw markdown from editor or AI-generated output
raw_md = """---
status: draft
author: loopx-bot
---
# New Feature: Health Packet Streaming
This implementation adds real-time health metrics...
"""
# Parse into typed schema
item = ContentItem.from_markdown(raw_md)
# ContentItem validates frontmatter, extracts body, computes content hash
# Execute full lifecycle pipeline
lifecycle = ItemLifecycle()
result = lifecycle.process(item)
# Internally: validation → transition → persistence → projection trigger
# Inspect attached evidence
print(result.evidence.signature) # Cryptographic content hash
print(result.evidence.lineage) # Provenance chain
The process() method encapsulates all six pipeline stages, returning a structured result with full auditability.
Multi-Step Pipeline: Issue-Fix Acceptance Loop
Complex data processing requires orchestrated multi-step pipelines. The Issue-Fix capability ([loopx/capabilities/issue_fix/acceptance_loop.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/issue_fix/acceptance_loop.py)) implements this pattern:
from loopx.capabilities.issue_fix.acceptance_loop import AcceptanceLoop
loop = AcceptanceLoop()
# Stage: Pre-flight validation
preflight = loop.validate_patch(patch_data)
if not preflight.passed:
return loop.reject(preflight.errors)
# Stage: Automated review
review = loop.analyze_impact(preflight.normalized_patch)
# Stage: Human or delegated gate
gate = loop.request_approval(review.assessment)
if gate.approved:
# Stage: Commit with full evidence chain
commit = loop.commit(gate.stamped_patch)
return loop.project_to_dashboard(commit.evidence)
Each stage generates its own transition, enabling partial rollback and fine-grained observability.
Summary
LoopX data processing is built on three architectural principles that emerge from the source code:
- Capability ownership: Each domain (reward memory, content ops, issue fix) controls its own ingestion, validation, and transition logic while conforming to kernel contracts
- Kernel centrality: All durable state flows through a single, typed, auditable store with atomic persistence
- Observability by design: Health packets and evidence chains are first-class outputs, not afterthoughts
The result is a provider-neutral, reproducible, and debuggable data processing foundation for long-running AI workflows.
Frequently Asked Questions
What is the role of the Capability Registry in LoopX data processing?
The Capability Registry ([loopx/capabilities/registry.py](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py)) maps capability identifiers to their execution handlers. It serves as the dispatch layer that routes validated transitions to the correct state mutation logic, ensuring type-safe, capability-isolated data processing across the kernel.
How does LoopX maintain data consistency during provider interruptions?
LoopX uses atomic state transitions with dual-write persistence to .loopx/registry.json and project files. If a provider fails mid-execution, the kernel retains its last committed state; partial transitions are discarded. The next agent turn resumes from a known-good checkpoint with full evidence of prior operations.
Can LoopX data processing pipelines be extended with custom capabilities?
Yes. New capabilities implement the standard interface: ingestion parser, contract validator, transition handler, and optional health reporter. Register the capability in the registry with a unique ID, and it immediately participates in the canonical Agent → Capability → Provider → Kernel flow with full observability support.
How does LoopX separate sensitive data from public projections?
Capabilities enforce public-private boundaries at the validation stage. The kernel marks evidence with visibility scopes; projection sinks filter accordingly. For example, raw API keys ingested during reward memory processing remain in private evidence, while sanitized metrics appear in dashboard projections.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →