How to Resume a Broken cangjie-skill Pipeline from PIPELINESTATE.md
To resume a broken cangjie-skill pipeline, edit the PIPELINESTATE.md file in the repository root to reset the failed stage from ❌ to ⏳ (pending), then re-run python -m cangjie.pipeline to restart from that stage while automatically skipping previously completed work.
The kangarooking/cangjie-skill repository implements a seven-stage RIA-TV++ pipeline that transforms raw text into structured AI skills. When execution halts due to crashes, timeouts, or manual interruptions, the pipeline state persists in PIPELINESTATE.md, allowing you to resume cangjie-skill pipeline operations without reprocessing the Adler, Parallel Extract, or other successfully finished stages.
Understanding the Pipeline State File
PIPELINESTATE.md lives in the repository root (<repo-root>/PIPELINESTATE.md) and serves as the single source of truth for execution progress. This human-readable Markdown file contains a table tracking the seven methodology stages: Adler, Parallel Extract, Triple Verification, RIA++ Construction, Zettelkasten Linking, Pressure Test, and Delivery. Each row displays a status emoji: ✅ for completed, ❌ for error, ⏸ for paused, and ⏳ for pending.
Step-by-Step Recovery Process
Locate the State File
Navigate to the repository root and open PIPELINESTATE.md in any text editor. The file structure mirrors the seven-stage workflow defined in methodology/00-overview.md and the master specification in SKILL.md.
Identify the Failed Stage
Scan the Status column for the last ✅ symbol. The stage immediately below it bearing ❌ or ⏸ indicates where execution stopped. For example, if 2 – Parallel Extract shows ❌, the pipeline failed while running the five parallel extractors (framework, principle, case, counter-example, and glossary) defined in extractors/*.md.
Reset Stages to Pending
Change the status of the failed stage—and any subsequent stages—from ❌ or ⏸ to ⏳ (pending). Save the file. This modification signals the runner to re-execute these specific stages while preserving work already validated in earlier steps.
# PIPELINESTATE.md
| Stage | Status |
|-------|--------|
| 1 – Adler | ✅ |
| 2 – Parallel Extract | ⏳ | <!-- changed from ❌ -->
| 3 – Triple Verification | ⏸ |
| 4 – RIA++ Construction | ⏸ |
| 5 – Zettelkasten Linking | ⏸ |
| 6 – Pressure Test | ⏸ |
| 7 – Delivery | ⏸ |
Clean Up Partial Outputs (Optional)
If the failed stage produced incomplete artifacts—such as fragmented SKILL.md files or corrupted entries in the output/ directory—delete or archive them before restarting. This prevents the extractor from detecting existing files and assuming the stage succeeded, which would mix corrupted data with fresh results.
Restart the Pipeline
Execute the pipeline command from the repository root:
python -m cangjie.pipeline
Alternatively, use the Make wrapper if configured in your environment:
make pipeline
The runner script reads PIPELINESTATE.md, detects the ⏳ status on the interrupted stage, and resumes execution from that point, skipping stages marked ✅.
Advanced Recovery Options
Force Start from a Specific Stage
To bypass automatic state detection and explicitly begin at a specific stage, use the --start-stage flag:
python -m cangjie.pipeline --start-stage 2
This command internally reads PIPELINESTATE.md but overrides the pending check to force execution to begin at stage 2 (Parallel Extract), regardless of the current status emoji.
Verification and Output
After the run completes, verify that PIPELINESTATE.md displays ✅ for all seven stages. Confirm successful regeneration by checking the output/ directory for BOOK_OVERVIEW.md, INDEX.md, and DIGEST.md, along with the individual skill modules generated by the templates in templates/.
Summary
- Single source of truth:
PIPELINESTATE.mdcontrols resumption; modify only the Status column, as stage names and order are fixed by the RIA-TV++ methodology. - Safe to retry: Stages marked ✅ are automatically skipped during restart, preventing redundant processing.
- Manual reset required: Change failed stages from ❌ to ⏳ to trigger re-execution.
- Cleanup recommended: Remove partial
SKILL.mdfragments or broken files from theoutput/directory to avoid data corruption. - Entry points: Use
python -m cangjie.pipelineormake pipelineto resume; add--start-stage Nto force a specific starting point.
Frequently Asked Questions
What happens if I don't clean up partial outputs from a failed stage?
The extractor scripts may detect existing files and assume the stage completed successfully, causing the pipeline to skip necessary work. This risks mixing corrupted data with fresh results, so always remove incomplete artifacts from the output/ folder or */SKILL.md paths before resuming.
Can I resume the pipeline if I accidentally deleted PIPELINESTATE.md?
If the state file is missing, the pipeline assumes a fresh start and will execute all seven stages from the beginning. Without PIPELINESTATE.md, the system cannot determine which steps completed successfully, so you must reprocess the entire workflow.
Does the --start-stage flag permanently modify PIPELINESTATE.md?
No, the --start-stage flag only affects the current execution session. While the pipeline reads the state file internally to locate resources, you must manually edit PIPELINESTATE.md to permanently change stage statuses for subsequent runs.
Where can I find logs to debug why a stage keeps failing?
Inspect the error messages recorded in PIPELINESTATE.md and check the log/ directory or stdout output from the verification scripts. The methodology documentation in methodology/00-overview.md explains each stage's validation criteria, helping you identify why the Triple Verification or Pressure Test stages repeatedly error out.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →