# How Breakpoint/Resume Works with PIPELINE_STATE.md in cangjie-skill

> Learn how cangjie-skill uses PIPELINE_STATE.md to enable fault-tolerant execution and resume pipelines automatically from the last successful checkpoint after interruptions.

- Repository: [kangarooking/cangjie-skill](https://github.com/kangarooking/cangjie-skill)
- Tags: internals
- Published: 2026-07-22

---

**The cangjie-skill pipeline implements fault-tolerant execution by storing stage completion status in `books/<book-slug>/PIPELINE_STATE.md`, enabling automatic resumption from the last successful checkpoint after interruptions.**

The kangarooking/cangjie-skill repository provides a strictly ordered, five-stage workflow that converts long-form content into executable Claude skills. Because this transformation process can be lengthy and subject to timeouts or crashes, the system implements **breakpoint/resume functionality** through a simple Markdown checkpoint file. The [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) file serves as the single source of truth for pipeline state, eliminating the need for external databases while maintaining full auditability.

## What is PIPELINE_STATE.md?

The checkpoint file resides at `books/<book-slug>/PIPELINE_STATE.md` inside each book's output folder. It records the **current execution stage**, the **set of completed artifacts**, and the **status of each skill** (e.g., "in-progress", "done", "failed").

The format follows a human-readable Markdown checklist:

```markdown
- [x] Stage 0 – 整书理解 (DONE)
- [ ] Stage 1 – 并行提取
- [ ] Stage 2 – 技能生成

```

According to line 54 of [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md), this file functions as "流水线状态: 当前阶段 + 各 skill 进度 (断点续跑用)"—a running log of current phase and skill progress designed specifically for breakpoint/resume scenarios.

## How Breakpoint/Resume Works

### Startup State Detection

When the pipeline initializes, it first checks for the existence of [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) in the book's output directory.

- **If the file exists**: The runner reads the recorded stage and jumps directly to the next unfinished stage, avoiding redundant reprocessing.
- **If the file does not exist**: The pipeline assumes a fresh run and creates the file with Stage 0 marked as "in-progress".

As documented at line 72 of [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md): "断点续跑: 开始前先检查 `books/<slug>/PIPELINE_STATE.md` 是否存在。存在则读取并从记录的阶段续跑,不要从头重来."

### Stage Completion Tracking

After each stage finishes—such as producing [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md) in Stage 0—the orchestrator updates the checkpoint file. The script locates the line corresponding to the completed stage and toggles the checkbox from `[ ]` to `[x]`, appending a completion status marker.

### Graceful Interruption Handling

If the process terminates unexpectedly (Ctrl-C, system crash, or cloud timeout), the [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) remains on disk with the last successfully completed stage still marked. Because the file uses plain Markdown, no database locks or cleanup transactions are required.

### Resume Execution Logic

On subsequent runs, the orchestrator re-parses the checkpoint file to determine the **last completed stage index**. It then continues execution from the following stage onward, skipping expensive operations like parallel extraction or pressure testing that would otherwise repeat work.

### Pipeline Finalization

When Stage 5 completes and generates [`DIGEST.md`](https://github.com/kangarooking/cangjie-skill/blob/main/DIGEST.md), the system marks the entire pipeline as finished:

```markdown
- [x] Stage 5 – 交付 (DONE)
- [x] Pipeline complete

```

At this point, all artifacts are materialized and the checkpoint file can be archived or deleted.

## Implementation Example

The following Python logic mirrors the checkpoint mechanism described in the skill documentation:

```python
import pathlib
import markdown

PIPELINE_PATH = pathlib.Path("books") / slug / "PIPELINE_STATE.md"

def load_state():
    """Read the checkpoint file and return the last completed stage."""
    if not PIPELINE_PATH.exists():
        # Fresh run – create initial state

        PIPELINE_PATH.write_text("- [ ] Stage 0 – 整书理解\n- [ ] Stage 1 – 并行提取\n…\n")
        return -1        # No stage completed yet

    # Parse Markdown checklist

    lines = PIPELINE_PATH.read_text().splitlines()
    for i, line in enumerate(lines):
        if line.startswith("- [ ]"):
            return i - 1   # The previous stage is the last completed one

    return len(lines) - 1  # All stages done

def mark_stage_done(stage_idx, description):
    """Mark a stage as completed in the checkpoint file."""
    lines = PIPELINE_PATH.read_text().splitlines()
    # Replace the line for this stage with a checked box

    for i, line in enumerate(lines):
        if line.startswith(f"- [ ] Stage {stage_idx}"):
            lines[i] = f"- [x] Stage {stage_idx} – {description} (DONE)"
            break
    PIPELINE_PATH.write_text("\n".join(lines))

# Example usage:

last_done = load_state()
if last_done < 0:
    run_stage_0()
    mark_stage_done(0, "整书理解")
if last_done < 1:
    run_stage_1()
    mark_stage_done(1, "并行提取")

# … continue through stages 2‑5 …

```

This implementation demonstrates how the pipeline uses simple file I/O to achieve durable state management without external dependencies, as referenced in lines 70–73 of [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md).

## Design Benefits

The checkpoint architecture provides four key advantages:

- **Simplicity** – A human-readable Markdown checklist eliminates the need for complex state databases or transactional systems.
- **Transparency** – Users can open [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) at any time to inspect exactly where the pipeline stopped and which artifacts are ready.
- **Idempotence** – Each stage is designed to be safely rerunnable; the checkpoint merely directs the orchestrator where to start without risking duplicate outputs.
- **Auditability** – Because every intermediate artifact remains on disk alongside the checkpoint index, the file provides a concise, reproducible record of progress.

## Summary

- **PIPELINE_STATE.md** lives at `books/<book-slug>/PIPELINE_STATE.md` and stores the pipeline's execution state as a Markdown checklist.
- **Breakpoint/resume functionality** works by checking for this file at startup; if present, the pipeline resumes from the next incomplete stage.
- **Stage completion** updates the file by marking checkboxes as `[x]`, creating a durable record that survives crashes and timeouts.
- **Implementation** relies on simple file I/O and string parsing rather than external databases, making the system portable and transparent.

## Frequently Asked Questions

### What happens if I delete PIPELINE_STATE.md mid-run?

If the checkpoint file is deleted, the pipeline treats the next execution as a fresh start beginning at Stage 0. It will recreate the file and regenerate artifacts from the beginning. To resume from a specific point, ensure the file remains intact with the correct stage marked as completed.

### Can I manually edit the checkpoint file to skip stages?

Yes. Because [`PIPELINE_STATE.md`](https://github.com/kangarooking/cangjie-skill/blob/main/PIPELINE_STATE.md) uses plain Markdown, you can manually check boxes to mark stages as complete. However, this is only safe if the corresponding output artifacts (e.g., [`BOOK_OVERVIEW.md`](https://github.com/kangarooking/cangjie-skill/blob/main/BOOK_OVERVIEW.md), skill directories) actually exist in the `books/<slug>/` folder, as subsequent stages depend on these files being present.

### How does the pipeline handle failed stages?

The checkpoint file marks stages as completed only after successful execution. If a stage fails, the corresponding checkbox remains `[ ]`, and the next run will retry that same stage. This prevents the pipeline from advancing with partial or corrupted intermediate data.

### Is PIPELINE_STATE.md the only state storage mechanism?

According to the cangjie-skill architecture documented in [`SKILL.md`](https://github.com/kangarooking/cangjie-skill/blob/main/SKILL.md) and [`methodology/00-overview.md`](https://github.com/kangarooking/cangjie-skill/blob/main/methodology/00-overview.md), yes. The pipeline relies solely on this file and the presence of intermediate artifacts on disk. There is no external database, registry, or cloud state store required, making the system fully portable and easy to inspect.