# How LoopX Handles Data Processing: From Ingestion to Kernel State

> Discover how LoopX processes data from ingestion to kernel state via a deterministic control-plane pipeline. Learn about observation validation, state transitions, and durable kernel persistence.

- Repository: [huangruiteng/loopx](https://github.com/huangruiteng/loopx)
- Tags: how-to-guide
- Published: 2026-08-14

---

**LoopX processes data through a deterministic control-plane pipeline that ingests raw observations, validates them against typed contracts, transforms them into state transitions, and persists results to a durable kernel with built-in health monitoring and projection capabilities.**

LoopX is an open-source control-plane runtime for AI workflows. Understanding how it handles data processing reveals a carefully architected separation between *what* executes (capabilities), *who* executes it (providers), and *where* state lives (the kernel). This article examines the actual source code implementation to show how data flows from raw input to durable, observable output.

---

## Data Processing Architecture in LoopX

The LoopX data processing pipeline follows a canonical flow documented in [[`docs/architecture.md`](https://github.com/huangruiteng/loopx/blob/main/docs/architecture.md)](https://github.com/huangruiteng/loopx/blob/main/docs/architecture.md):

```

Agent → Capability → Provider → Kernel → Provider readback → Capability transition → Kernel

```

This bidirectional flow ensures that every data operation leaves an auditable trace in kernel state while remaining agnostic to the underlying AI provider (Codex App, Claude Code, or custom runners).

### Core Components

| Component | Responsibility | Key Source File |
|-----------|---------------|---------------|
| **Capability** | Defines *what* processing happens; owns ingestion, validation, and transition logic | [[`loopx/capabilities/registry.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py) |
| **Provider** | Executes tooling calls; returns raw outputs for capability consumption | Provider-specific adapters (Codex, Claude, shell) |
| **Kernel** | Single source of truth for durable state (`todos`, `gates`, `evidence`, `quota`) | [[`loopx/state/model.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/state/model.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/state/model.py) |
| **Health Module** | Emits observable metrics for monitoring and debugging | Capability-specific health modules |
| **Projection** | Renders kernel state to human-facing surfaces (dashboards, Kanban boards) | [[`apps/presentation/dashboard/README.md`](https://github.com/huangruiteng/loopx/blob/main/apps/presentation/dashboard/README.md)](https://github.com/huangruiteng/loopx/blob/main/apps/presentation/dashboard/README.md) |

---

## The Six-Stage Data Processing Pipeline

### Stage 1: Ingestion

Every capability implements its own ingestion logic to parse raw provider output. The **Reward Memory** capability demonstrates this pattern in [[`loopx/capabilities/reward_memory/ingestion.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/ingestion.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/ingestion.py):

```python
from loopx.capabilities.reward_memory.ingestion import ingest_memory

# Raw payload from external source (e.g., OpenViking corpus)

raw_payload = {
    "corpus_id": "proj-123",
    "documents": [
        {"id": "doc-1", "content": "Implementation notes...", "metadata": {"version": "2.1"}}
    ]
}

# Parse, enrich with metadata, and prepare for validation

ingest_result = ingest_memory(raw_payload)

```

The ingestion layer handles:
- Schema-compatible parsing
- Metadata extraction
- Initial error classification

### Stage 2: Validation

Capabilities enforce **typed contracts** and **public-private boundaries** before any state mutation. The Reward Memory contract ([[`loopx/capabilities/reward_memory/contract.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/contract.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/contract.py)) validates that:

- Required fields are present and correctly typed
- Evidence-gate rules are satisfied
- No private data leaks across capability boundaries

```python
if ingest_result.is_valid:
    # Proceed to kernel transition

    loopx_kernel.apply_transition(ingest_result.transition)
else:
    # Validation failures become structured error evidence

    loopx_kernel.record_evidence(ingest_result.validation_errors)

```

### Stage 3: Transition

The kernel applies **deterministic state transitions** through its registry. In [[`loopx/capabilities/registry.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py), capabilities register their transition handlers:

```python

# Simplified registry pattern

class CapabilityRegistry:
    def register(self, capability_id: str, handler: TransitionHandler):
        self._handlers[capability_id] = handler
    
    def apply_transition(self, capability_id: str, transition: StateTransition):
        handler = self._handlers[capability_id]
        updated_state = handler.execute(self._kernel_state, transition)
        self._persist(updated_state)

```

Transitions operate on four core state primitives defined in [[`loopx/state/model.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/state/model.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/state/model.py):
- **`todos`**: Actionable work items
- **`gates`**: Decision checkpoints requiring human or automated approval
- **`evidence`**: Immutable records of what occurred
- **`quota`**: Resource consumption tracking

### Stage 4: Health Monitoring

Every capability emits **health packets** for observability. The Reward Memory health module ([[`loopx/capabilities/reward_memory/health.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/health.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/reward_memory/health.py)) generates structured diagnostics:

```python
from loopx.capabilities.reward_memory.health import health_report

metrics = health_report(ingest_result)
print(metrics)

# Output:

# {

#   'pipeline': {

#     'corpus_present': True,

#     'index_present': True,

#     'document_count': 147,

#     'last_ingestion_at': '2025-01-15T09:23:11Z'

#   }

# }

```

Health checks enable:
- Real-time pipeline monitoring
- Automated alerting on data quality degradation
- Debugging support for multi-turn AI workflows

### Stage 5: Persistence

The kernel serializes state to **[`.loopx/registry.json`](https://github.com/huangruiteng/loopx/blob/main/.loopx/registry.json)** and optional project-scoped files. This dual-write pattern ensures:

- **Global consistency**: The registry maintains cross-project references
- **Local reproducibility**: Project files enable standalone execution

Persistence is atomic: transitions either fully commit or roll back, preventing partial state corruption during provider interruptions.

### Stage 6: Projection

Processed data becomes human-readable through **projection sinks**. The dashboard application ([[`apps/presentation/dashboard/README.md`](https://github.com/huangruiteng/loopx/blob/main/apps/presentation/dashboard/README.md)](https://github.com/huangruiteng/loopx/blob/main/apps/presentation/dashboard/README.md)) renders kernel state as:

- Interactive Kanban boards (Lark integration)
- HTML reports with embedded evidence
- Real-time operational metrics

Projection decouples data processing from presentation, allowing the same kernel state to serve multiple audience needs.

---

## Complete Example: Content-Ops Item Lifecycle

The **Content Ops** capability demonstrates the full pipeline for markdown content processing in [[`loopx/capabilities/content_ops/item_lifecycle.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/content_ops/item_lifecycle.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/content_ops/item_lifecycle.py):

```python
from loopx.capabilities.content_ops.item_lifecycle import ItemLifecycle
from loopx.capabilities.content_ops.schemas import ContentItem

# Raw markdown from editor or AI-generated output

raw_md = """---
status: draft
author: loopx-bot
---

# New Feature: Health Packet Streaming

This implementation adds real-time health metrics...
"""

# Parse into typed schema

item = ContentItem.from_markdown(raw_md)

# ContentItem validates frontmatter, extracts body, computes content hash

# Execute full lifecycle pipeline

lifecycle = ItemLifecycle()
result = lifecycle.process(item)

# Internally: validation → transition → persistence → projection trigger

# Inspect attached evidence

print(result.evidence.signature)  # Cryptographic content hash

print(result.evidence.lineage)    # Provenance chain

```

The `process()` method encapsulates all six pipeline stages, returning a structured result with full auditability.

---

## Multi-Step Pipeline: Issue-Fix Acceptance Loop

Complex data processing requires **orchestrated multi-step pipelines**. The Issue-Fix capability ([[`loopx/capabilities/issue_fix/acceptance_loop.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/issue_fix/acceptance_loop.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/issue_fix/acceptance_loop.py)) implements this pattern:

```python
from loopx.capabilities.issue_fix.acceptance_loop import AcceptanceLoop

loop = AcceptanceLoop()

# Stage: Pre-flight validation

preflight = loop.validate_patch(patch_data)
if not preflight.passed:
    return loop.reject(preflight.errors)

# Stage: Automated review

review = loop.analyze_impact(preflight.normalized_patch)

# Stage: Human or delegated gate

gate = loop.request_approval(review.assessment)
if gate.approved:
    # Stage: Commit with full evidence chain

    commit = loop.commit(gate.stamped_patch)
    return loop.project_to_dashboard(commit.evidence)

```

Each stage generates its own transition, enabling partial rollback and fine-grained observability.

---

## Summary

LoopX data processing is built on three architectural principles that emerge from the source code:

- **Capability ownership**: Each domain (reward memory, content ops, issue fix) controls its own ingestion, validation, and transition logic while conforming to kernel contracts
- **Kernel centrality**: All durable state flows through a single, typed, auditable store with atomic persistence
- **Observability by design**: Health packets and evidence chains are first-class outputs, not afterthoughts

The result is a **provider-neutral, reproducible, and debuggable** data processing foundation for long-running AI workflows.

---

## Frequently Asked Questions

### What is the role of the Capability Registry in LoopX data processing?

The **Capability Registry** ([[`loopx/capabilities/registry.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py)](https://github.com/huangruiteng/loopx/blob/main/loopx/capabilities/registry.py)) maps capability identifiers to their execution handlers. It serves as the dispatch layer that routes validated transitions to the correct state mutation logic, ensuring type-safe, capability-isolated data processing across the kernel.

### How does LoopX maintain data consistency during provider interruptions?

LoopX uses **atomic state transitions** with dual-write persistence to [`.loopx/registry.json`](https://github.com/huangruiteng/loopx/blob/main/.loopx/registry.json) and project files. If a provider fails mid-execution, the kernel retains its last committed state; partial transitions are discarded. The next agent turn resumes from a known-good checkpoint with full evidence of prior operations.

### Can LoopX data processing pipelines be extended with custom capabilities?

Yes. New capabilities implement the standard interface: ingestion parser, contract validator, transition handler, and optional health reporter. Register the capability in the registry with a unique ID, and it immediately participates in the canonical `Agent → Capability → Provider → Kernel` flow with full observability support.

### How does LoopX separate sensitive data from public projections?

Capabilities enforce **public-private boundaries** at the validation stage. The kernel marks evidence with visibility scopes; projection sinks filter accordingly. For example, raw API keys ingested during reward memory processing remain in private evidence, while sanitized metrics appear in dashboard projections.