How Checkpoint and Resume Works with PIPELINE_STATE.md in Cangjie-Skill

The cangjie-skill pipeline uses a markdown checklist in PIPELINE_STATE.md to track which of its seven stages have completed, allowing interrupted runs to resume exactly where they left off.

The cangjie-skill repository implements a robust checkpoint and resume system for its RIA-TV++ data processing pipeline. According to the source code in [SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), the pipeline tracks execution state through a human-readable markdown file rather than hidden binary logs or complex databases.

The Seven-Stage Pipeline Structure

The RIA-TV++ pipeline processes data through seven sequential stages:

  1. Stage 0 – Adler
  2. Stage 1 – Parallel Extract
  3. Stage 2 – RIA-Plus
  4. Stage 3 – Parallel Analysis
  5. Stage 4 – Data Fusion
  6. Stage 5 – Report Generation
  7. Stage 6 – Final Output

Each stage can take significant time to complete. The checkpoint system ensures that if a run fails during Stage 4, you do not need to re-run Stages 0-3.

How the Checkpoint Mechanism Works

Writing Progress to PIPELINE_STATE.md

After each stage finishes successfully, the pipeline updates [PIPELINE_STATE.md](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) to mark that stage as complete. The file uses a simple markdown checklist format:

- [x] Stage 0 – Adler
- [x] Stage 1 – Parallel Extract
- [ ] Stage 2 – RIA-Plus
- [ ] Stage 3 – Parallel Analysis
- [ ] Stage 4 – Data Fusion
- [ ] Stage 5 – Report Generation
- [ ] Stage 6 – Final Output

The notation follows standard markdown task list syntax:

  • [ ] (unchecked) — stage not yet run or was reset
  • [x] (checked) — stage completed successfully

The driver code in SKILL.md performs the actual file update using pattern replacement on the markdown text.

How the Resume Mechanism Works

Parsing State on Startup

When the pipeline starts, it reads PIPELINE_STATE.md if it exists, or creates it with all stages unchecked on first run. The resume logic parses each line with a regular expression to determine which stages to skip:

import pathlib
import re

STATE_FILE = pathlib.Path('PIPELINE_STATE.md')

# Parse current state from markdown checklist

state_text = STATE_FILE.read_text()
state = {
    m.group(2): m.group(1) == 'x'
    for m in re.finditer(r'- \[(.)\] (.+)', state_text)
}

The state dictionary maps stage names to boolean completion status.

Conditional Stage Execution

The driver iterates through the ordered stage list and executes only unchecked stages:

stages = [
    ('Stage 0 – Adler', run_stage0),
    ('Stage 1 – Parallel Extract', run_stage1),
    ('Stage 2 – RIA-Plus', run_stage2),
    # ... stages 3-6

]

for name, func in stages:
    if not state.get(name, False):
        func()  # Execute stage

        
        # Update checkpoint: mark as complete

        text = STATE_FILE.read_text()
        updated = re.sub(
            rf'(\- \[ \] {re.escape(name)})',
            f'- [x] {name}',
            text
        )
        STATE_FILE.write_text(updated)

This ensures idempotent resume behavior: re-running the pipeline after any interruption automatically skips completed work.

Manual Control and Debugging

Because PIPELINE_STATE.md is plain text, you can manipulate it directly without special tools.

Resetting a Specific Stage

To force re-run of Stage 2 for debugging:


# Edit PIPELINE_STATE.md

# Change:  - [x] Stage 2 – RIA-Plus

# To:      - [ ] Stage 2 – RIA-Plus

The next pipeline run will execute Stage 2 and all subsequent stages.

Complete Reset

To restart from scratch:


# Delete or truncate the state file

rm PIPELINE_STATE.md

The pipeline will recreate it with all stages unchecked.

Key Implementation Files

File Purpose
[PIPELINE_STATE.md](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) Persistent checkpoint storage using markdown checklist format
[SKILL.md](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) Main driver script containing checkpoint read/write logic
[README.en.md](https://github.com/kangarooking/cangjie-skill/blob/main/README.en.md) Documentation of the seven-stage RIA-TV++ pipeline
Stage documentation files (01-stage0-adler.md, etc.) Implementation details for each checkpointed stage

Design Advantages

The markdown-based approach provides several benefits over traditional checkpoint systems:

  • Visibility — Progress is immediately readable without special tools
  • Version control friendly — Git diffs show exactly which stages completed
  • Human editable — Developers can manually adjust state when needed
  • No database dependencies — Single file, no external services required
  • Cross-platform — Works identically on all operating systems

Summary

  • Checkpoint mechanism writes stage completion status to PIPELINE_STATE.md using standard markdown checkboxes ([x] vs [ ])
  • Resume functionality parses the markdown file on startup and skips all stages marked complete, continuing from the first unchecked item
  • Implementation resides in SKILL.md using regex-based text manipulation for state reading and updating
  • Manual control is fully supported through direct file editing, enabling debugging and selective re-runs
  • Seven stages (Adler through Final Output) are tracked sequentially in the RIA-TV++ pipeline

Frequently Asked Questions

What happens if I delete PIPELINE_STATE.md?

The pipeline creates a fresh state file with all seven stages unchecked on the next run. This effectively resets progress and causes a full re-execution of the entire pipeline.

Can I run individual stages without modifying the state file?

No, the current implementation in SKILL.md does not support stage selection via command-line arguments. You must manually edit PIPELINE_STATE.md to uncheck specific stages, or implement a wrapper that injects stage selection before the main driver loop.

Is the state file safe for concurrent access?

The implementation uses simple file read/write operations without file locking. Concurrent pipeline runs could potentially corrupt PIPELINE_STATE.md if they write simultaneously. The design assumes single-process execution.

Why markdown instead of JSON or YAML?

Markdown checklist format prioritizes human readability and editability. JSON or YAML would require parsing libraries and obscure the state behind syntax. The regex-based parsing in SKILL.md is sufficient for the fixed seven-stage structure and maintains the file's purpose as both machine-readable checkpoint and human-readable progress indicator.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →