How to Use LoopX for Data Processing: A Complete Guide to Provider-Neutral Pipelines

LoopX is a lightweight, provider-neutral control-plane that orchestrates long-running data processing pipelines through value-connectors, quota-based scheduling, and auditable state management, requiring only Python standard library dependencies.

LoopX provides a lightweight, provider-neutral control-plane that lets you run long-running, multi-step data-processing pipelines while keeping every turn's objective, evidence, and hand-off visible and auditable. Unlike monolithic data orchestrators, LoopX separates the execution environment from state management, allowing you to connect external workers, LLMs, or custom scripts while maintaining an immutable record of decisions. This guide walks through the architecture, installation, and practical implementation of data processing workflows using the LoopX framework.

Understanding LoopX's Five-Layer Architecture for Data Processing

LoopX organizes data processing workflows into five distinct layers that interact through well-defined contracts. Each layer serves a specific function in maintaining auditable, human-in-the-loop pipelines.

Goal and State Layer

The Goal and State layer stores the active objective, scoped gates, todos, evidence, and quota. According to docs/state-interaction-model.md, this layer acts as the single source of truth for your pipeline's current condition, tracking every piece of evidence generated during data extraction and transformation.

Quota Layer

The Quota layer decides whether a registered agent may act now, what kind of work it may do, and when to pause for human judgment. As documented in docs/quota-allocation.md, this prevents runaway resource consumption by requiring explicit slots before executing expensive data operations.

Agent Runtime Bridge

The Agent Runtime Bridge connects external execution environments—whether OpenAI Codex, Anthropic Claude, or custom Python workers—to the LoopX kernel. The implementation in loopx/worker_bridge.py handles authentication, request serialization, and result validation between your data processing scripts and the core state machine.

Operator Surface

The Operator Surface renders read-first dashboards (web UI or Lark Kanban) that display current goals, pending todos, and evidence without being the source of truth. The dashboard implementation in apps/presentation/dashboard/README.md allows stakeholders to review extracted data without modifying the underlying state.

Capabilities Layer

Capabilities encapsulate reusable, typed work-lanes that consume provider output, validate it, and propose state transitions. For data processing, you typically use value-connector capabilities, documented in docs/capabilities/value-connectors/README.md, which pull raw data from external sources, normalize it, and write structured results back as evidence.

Installing LoopX and Connecting Your Project

LoopX requires no repository clone and runs with Python standard library dependencies only. Install the control-plane using the official installation script.


# Install LoopX (single-line script)

curl -fsSL https://raw.githubusercontent.com/huangruiteng/loopx/main/scripts/install-from-github.sh | bash
export PATH="$HOME/.local/bin:$PATH"

# Verify installation

loopx doctor

After installation, connect LoopX to your project directory. This creates a hidden .loopx/ directory containing the kernel state and registry.

cd <project-dir>
loopx connect  # Creates .loopx/registry.json

Building a Data Processing Pipeline with Value Connectors

Value-connectors are specialized capabilities that handle data ingestion and transformation. They fetch raw data from external APIs or files, normalize the content, and write structured evidence back to the LoopX state.

A typical data processing workflow follows these steps:

  1. Define the goal using the guided wizard to establish the extraction objective.
  2. Run a value-connector to fetch and process data from external sources.
  3. Inspect evidence through the operator dashboard to verify results.
  4. Resolve gates if human review blocks automatic progression.
  5. Consume quota slots to trigger downstream aggregation capabilities.

Execute this workflow using the LoopX CLI:


# Create a new goal (guided wizard)

loopx start-goal --guided --project . \
  --goal-text "Extract quarterly financial metrics from SEC filings"

# Run the finance value-connector capability

loopx value-connectors finance-discovery \
  --source-url https://www.sec.gov/Archives/edgar/data/...

# Check status and view evidence rows

loopx status

# If a gate blocks progress, claim and resolve the todo

loopx todo claim

# ...perform human review...

loopx todo update  # Mark gate as resolved

# Let quota schedule the next processing slice

loopx quota should-run
loopx quota spend-slot

All CLI commands are implemented in loopx/cli.py, which provides the interface for status checking, quota management, and todo resolution.

Running the Finance Value Discovery Example

The packages/loopx-finance-value-discovery package demonstrates a complete data extraction pipeline. This connector fetches public market data, creates financial-metric evidence rows, and returns JSON payloads that LoopX writes to the kernel.

Run the example directly without configuring a full goal:

python -m packages.loopx_finance_value_discovery.example \
  --ticker AAPL \
  --quarter 2023Q4

This script, located at packages/loopx-finance-value-discovery/example.py, serves as a reference implementation for building custom connectors that interface with the LoopX state model.

Managing Quota and Human-in-the-Loop Gates

The quota system ensures data processing pipelines remain cost-controlled and auditable. Before executing any capability, LoopX checks docs/quota-allocation.md contracts to determine if the agent has permission to proceed.

When a gate blocks progress—such as when confidence thresholds require human validation—the pipeline pauses and creates a todo item. Operators claim these todos through the CLI, perform necessary review (examining the compact evidence stored in the state model), and mark gates as resolved before quota allows the next execution slice.

This architecture guarantees that all intermediate artifacts remain stored as compact evidence, visible in the state diagram's evidence column, while preventing unauthorized or excessive data processing operations.

Summary

  • LoopX provides a provider-neutral control-plane for long-running data processing pipelines, separating execution environments from state management.
  • The architecture consists of five layers: Goal and State, Quota, Agent Runtime Bridge, Operator Surface, and Capabilities, interacting through contracts defined in docs/state-interaction-model.md and docs/quota-allocation.md.
  • Value-connectors handle data ingestion by fetching external sources, normalizing data, and writing structured evidence back to the kernel.
  • The loopx CLI in loopx/cli.py provides commands for goal creation (start-goal), connector execution (value-connectors), and quota management (quota should-run, quota spend-slot).
  • All operations require explicit quota slots and support human-in-the-loop gates to maintain auditability and prevent resource overconsumption.

Frequently Asked Questions

What makes LoopX different from Apache Airflow or Prefect for data processing?

LoopX separates the execution environment from the control-plane, allowing you to run data processing scripts in any runtime (local Python, Codex, Claude) while maintaining an auditable state of evidence and decisions. Unlike Airflow's DAG-centric model, LoopX uses goal-oriented state machines with quota-based scheduling and explicit human gates, making it ideal for workflows requiring frequent human judgment or multi-provider LLM orchestration.

Do I need to install Python dependencies beyond the standard library to use LoopX?

No. LoopX requires only the Python standard library for its core operation. External data processing dependencies (such as pandas, requests, or numpy) are managed within your specific value-connector packages or execution environments, not by the LoopX kernel itself. The installation script at scripts/install-from-github.sh sets up the CLI without additional runtime dependencies.

How does LoopX handle authentication for external data sources in value-connectors?

Authentication credentials are managed within the specific value-connector implementation or through environment variables accessible to the Agent Runtime Bridge. The loopx/worker_bridge.py module passes context to external workers without storing sensitive tokens in the LoopX state, ensuring that credentials remain in the execution environment while processing results are written back as evidence.

Can I integrate LoopX with existing Lark or Slack workflows for data approval?

Yes. The Operator Surface layer supports multiple presentation formats, including Lark Kanban integration documented in apps/presentation/dashboard/README.md. When a gate requires human review, LoopX can notify operators through these channels, allowing stakeholders to approve data processing steps directly from their existing collaboration tools while the kernel maintains the authoritative state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →