# How Peer Agents Use Todo Claim and Todo Update to Coordinate Parallel Work

> Discover how LoopX peer agents use todo claim and todo update to coordinate parallel work, ensuring race-condition-free execution via centralized state validation.

- Repository: [huangruiteng/loopx](https://github.com/huangruiteng/loopx)
- Tags: deep-dive
- Published: 2026-09-02

---

**Peer agents in LoopX use atomic `todo claim` operations to acquire exclusive locks on work items and `todo update` to modify them, ensuring race-condition-free parallel execution through centralized state validation.**

LoopX implements a robust peer-agent coordination system where multiple autonomous agents process shared goals without collision. By treating **todos** as the central unit of work in `loopx/control_plane/todos/`, the repository enforces strict claim-before-update semantics that prevent duplicate execution. This architecture allows peer agents to coordinate parallel work safely across distributed environments using a shared, auditable state store.

## The Coordination Architecture

### Centralized Todo Store and Exclusive Locking

LoopX stores all work items as todos in a central state accessible to every agent. To prevent race conditions when several agents target the same goal, the system requires an agent to **claim** a todo before modification. This claim acts as an exclusive lock recorded in the *claimed-by* metadata field, making ownership visible to all peers and ensuring that only one agent mutates the work item at a time.

### The Claim-Update-Release Lifecycle

Each todo progresses through a strict lifecycle enforced by validation middleware:

1. **Discovery** – Agents scan for open work matching their capabilities.
2. **Claim** – Exclusive lock acquisition with a unique agent identifier.
3. **Update** – Mutation of status, notes, or routing while holding the lock.
4. **Release** – Clearing the claim upon completion or supersession with a successor todo.

## Core Validation Components

### CLI Argument Validation ([`todo_argument_validation.py`](https://github.com/huangruiteng/loopx/blob/main/todo_argument_validation.py))

The `validate_todo_claim_options` and `validate_todo_update_options` functions in [`loopx/cli_commands/todo_argument_validation.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/cli_commands/todo_argument_validation.py) parse command-line input and enforce semantic rules before the system attempts persistence.

- For **claims** (lines 311–326), the validator mandates `--todo-id` and `--claimed-by` while rejecting unrelated flags that do not apply to the claim operation.
- For **updates** (lines 328–347), it ensures at least one mutable field is present and prevents logically conflicting flags, such as simultaneous `--claimed-by` and `--clear-claim` in the same command.

### Mutation Authority Enforcement ([`mutation_authority.py`](https://github.com/huangruiteng/loopx/blob/main/mutation_authority.py))

In [`loopx/control_plane/todos/mutation_authority.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/control_plane/todos/mutation_authority.py) (line 174), the system verifies that claims align with todo lifecycle constraints. A claim can only be placed on a todo with `open` status, and the claimant must be a registered peer. This prevents agents from seizing work already in progress or completed, acting as the primary gatekeeper against conflicting parallel access.

### Atomic Line Updates ([`line_update.py`](https://github.com/huangruiteng/loopx/blob/main/line_update.py))

The [`line_update.py`](https://github.com/huangruiteng/loopx/blob/main/line_update.py) module ([`loopx/control_plane/todos/line_update.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/control_plane/todos/line_update.py), line 304) applies persisted changes atomically. If a todo's status is not `open`, the claim operation aborts immediately. Failed transactions roll back without corrupting the claim state, maintaining lock integrity even during partial failures or system interruptions.

## The Four-Step Coordination Flow

### 1. Discovery

Agents begin by scanning the todo list using `loopx todo list --goal-id <id> --status open` to identify available work matching their capabilities.

### 2. Claim

Upon selecting work, an agent invokes the claim command with its identifier:

```bash
loopx todo claim \
    --goal-id my-goal \
    --todo-id todo_ab12cd34ef56 \
    --claimed-by codex-side-bypass

```

The claim commits only if [`mutation_authority.py`](https://github.com/huangruiteng/loopx/blob/main/mutation_authority.py) confirms the todo is unclaimed and open, writing the agent identifier to the *claimed-by* field. If another agent has already claimed the work, the validation fails and the command exits without modifying state.

### 3. Update

While holding the claim, the agent drives the todo through its lifecycle using update operations:

```bash
loopx todo update \
    --todo-id todo_ab12cd34ef56 \
    --status in_progress \
    --note "Processing batch #3"

```

`validate_todo_update_options` verifies that mutations are permitted for the current state, while [`line_update.py`](https://github.com/huangruiteng/loopx/blob/main/line_update.py) writes new values atomically. The system rejects updates that would violate the claim contract, such as attempts to modify the todo without holding the lock.

### 4. Release

Completion requires releasing the lock via `--clear-claim` or transitioning the todo to a completed state:

```bash
loopx todo update \
    --todo-id todo_ab12cd34ef56 \
    --clear-claim \
    --status completed \
    --note "All data processed"

```

Alternatively, agents can create a successor todo using `todo complete` or `todo supersede`, validated by `validate_successor_routing_options` to ensure proper routing chain continuity.

## Automatic Work Distribution via Heartbeat

The heartbeat agent in [`loopx/control_plane/heartbeat/agent.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/control_plane/heartbeat/agent.py) (lines 132–135) monitors for stale or unclaimed todos. When it detects work that should be picked up based on agent health and capability matching, it generates suggestions prompting specific peers to execute `todo claim`. This enables automatic load balancing and self-healing across the agent pool without centralized orchestration.

## Preventing Stale Locks and Duplicate Work

Because every claim binds to a concrete agent identifier, other peers observing the shared state automatically skip claimed todos. If an agent crashes, missing heartbeats trigger reclaim paths where surviving agents can seize orphaned work. The exclusive lock mechanism ensures that even under high parallelism, no two agents execute the same todo simultaneously.

## Summary

- LoopX uses **exclusive claims** as lightweight, auditable locks to prevent race conditions among peer agents.
- [`todo_argument_validation.py`](https://github.com/huangruiteng/loopx/blob/main/todo_argument_validation.py) enforces strict CLI rules for claim and update operations, rejecting malformed commands early.
- [`mutation_authority.py`](https://github.com/huangruiteng/loopx/blob/main/mutation_authority.py) validates lifecycle constraints at line 174, ensuring only `open` todos can be claimed by registered peers.
- [`line_update.py`](https://github.com/huangruiteng/loopx/blob/main/line_update.py) applies changes atomically at line 304 while preserving claim integrity during transaction failures.
- The **heartbeat agent** prompts peers to claim unclaimed work automatically, enabling resilient parallel coordination.

## Frequently Asked Questions

### What happens if two agents try to claim the same todo simultaneously?

The [`mutation_authority.py`](https://github.com/huangruiteng/loopx/blob/main/mutation_authority.py) layer acts as the gatekeeper. When two agents simultaneously attempt to claim an open todo, the first transaction to commit records its *claimed-by* identifier. The second request fails validation because the todo is no longer in `open` status for claiming, forcing the losing agent to select different work without creating duplicate execution.

### Can an agent update a todo without claiming it first?

No. While the CLI might superficially accept the command, the [`line_update.py`](https://github.com/huangruiteng/loopx/blob/main/line_update.py) persistence layer (line 304) and mutation authority enforce that modifications to sensitive fields require an active, matching claim. Updates without a valid *claimed-by* match are rejected to prevent unauthorized mutations during parallel execution.

### How does the system handle crashed agents that never release their claims?

The heartbeat mechanism in [`loopx/control_plane/heartbeat/agent.py`](https://github.com/huangruiteng/loopx/blob/main/loopx/control_plane/heartbeat/agent.py) monitors claim liveness. When an agent fails to check in, the heartbeat logic identifies stale claims and prompts healthy peers to execute `todo claim` with their own identifiers, effectively transferring ownership and allowing the work to proceed without manual administrator intervention.

### What is the difference between clearing a claim and superseding a todo?

**Clearing a claim** using `--clear-claim` removes the exclusive lock and updates the todo status, leaving the original work item intact in the history. **Superseding** creates a new successor todo using `todo supersede`, which links the completed work to a new work item and is validated by `validate_successor_routing_options` to ensure proper routing chain continuity. The former releases the lock on the current item, while the latter advances the workflow to the next phase.