How Playbook Distillation from Local Books Creates Experience Manuals Separate from the Case Vector Index
Playbook distillation processes local books through a three-stage pipeline that explicitly isolates the resulting experience manuals in oa/playbooks/ while returning into_case_index: false to prevent vectorization alongside case histories.
The handsomestWei/patent-disclosure-skill repository implements a specialized knowledge-management architecture where playbook distillation transforms static documents into reusable expertise without contaminating the similarity-searchable case library. This separation ensures that procedural guidance remains distinct from historical patent examination cases, allowing auditors to reference distilled strategies without conflating them with prior art vectors.
The Three-Step Distillation Workflow
The system orchestrates playbook creation through tools/oa/ingest_playbook.py, which coordinates pre-analysis, tooling verification, and final ingestion.
Step 1: Peek (Pre-Read) Analysis
Before committing resources to full distillation, the peek command performs lightweight keyword extraction to assess material relevance.
# Analyze a local PDF for patent examination keywords
python tools/oa/ingest_playbook.py peek --path my_book.pdf
In tools/oa/playbook.py, the peek function calls read_document to sample pages and returns a decision structure containing hint values ("likely", "unclear", etc.) based on keyword hits for terms like "审查答复", "创造性", and "新颖性". This gatekeeping prevents low-value books from entering the experience manual pipeline.
Step 2: Ensure Book-to-Skill Processor Availability
The workflow verifies that the external book-to-skill processor is installed and executable.
# Verify or install the external distillation processor
python tools/oa/ingest_playbook.py ensure-skill
This invokes logic in tools/oa/book_to_skill_setup.py to validate the environment before processing begins.
Step 3: Ingest Distilled Output
The final stage copies the processor's output into the isolated playbook directory.
# Ingest distilled knowledge into the playbook store
python tools/oa/ingest_playbook.py ingest \
--from-skill-dir ./distilled_output \
--source-path my_book.pdf \
--slug my-playbook
This command triggers ingest_distilled_skill in tools/oa/playbook.py, which orchestrates the _copy_distilled helper to transfer files and generates an _playbook.md index containing metadata fields like source_path, slug, and peek_decision.
Anatomy of an Experience Manual
Each distilled playbook resides in a dedicated subdirectory under oa/playbooks/{slug}, created via playbooks_root(oa_root) / slug in tools/oa/playbook.py. The directory structure includes:
- SKILL.md – Core procedural instructions for the patent domain
- cheatsheet.md – Quick-reference tactics for office action responses
- patterns.md – Recurring argument structures and templates
- _playbook.md – Machine-readable index linking the above components
These files comprise the experience manual, a standalone knowledge artifact distinct from the vectorized case history.
Isolating Playbooks from the Case Vector Index
The architectural separation between playbooks and searchable cases is enforced at multiple layers in the codebase.
Explicit Path Exclusion
In tools/oa/vault_layout.py, the list_playbook_index_paths helper specifically scans oa/playbooks/*/_playbook.md without adding those paths to the case ingestion queue. This ensures that when the vector index is built from case histories, playbook content remains excluded.
Indexing Metadata Flag
When ingest_distilled_skill completes, it returns a JSON payload containing "into_case_index": false, as implemented in tools/oa/playbook.py lines 32-34. This boolean flag signals upstream orchestrators that the material must not be embedded into the similarity search store used for retrieval-augmented generation.
Querying Experience Manuals at Runtime
During office action drafting, the system consults experience manuals through a separate retrieval path from case vectors. The runtime checks list_playbook_records to identify relevant manuals, then reads their cheatsheet.md or patterns.md contents.
Critically, these references carry the prefix 经验手册 (experience manual) rather than a numeric case_id, guaranteeing they remain outside the vector-search pipeline. This semantic distinction, documented in SKILL.md lines 212-214, ensures auditors receive procedural guidance without confusing prior case embeddings.
Summary
- Playbook distillation follows a three-stage workflow: peek analysis, tooling verification, and ingestion into
oa/playbooks/. - Experience manuals are structured as isolated directories containing SKILL.md, cheatsheet.md, and patterns.md, indexed by
_playbook.md. - Case vector separation is enforced by
list_playbook_index_pathsinvault_layout.pyand theinto_case_index: falsereturn value fromplaybook.py. - Runtime retrieval uses the
经验手册prefix to distinguish playbooks from case IDs, ensuring they bypass similarity search.
Frequently Asked Questions
What distinguishes a playbook from a case vector?
A playbook is a procedural experience manual containing distilled expertise (tactics, patterns, cheatsheets) stored in oa/playbooks/, while a case vector is an embedding of historical patent examination history used for similarity search. Playbooks are explicitly excluded from the vector index through path filtering in vault_layout.py.
How does the system prevent playbooks from entering the similarity search pipeline?
The ingest_distilled_skill function returns "into_case_index": false, and tools/oa/vault_layout.py isolates playbook paths from case ingestion logic. The list_playbook_index_paths function targets only oa/playbooks/*/_playbook.md, ensuring vector builders never process playbook content.
What content is generated during the distillation process?
The pipeline generates _playbook.md (metadata index), SKILL.md (domain procedures), cheatsheet.md (quick tactics), and patterns.md (reusable templates). These are copied from the external book-to-skill processor's output directory into the playbook structure via _copy_distilled in tools/oa/playbook.py.
How does the peek command determine if a book warrants distillation?
The peek command in tools/oa/ingest_playbook.py samples pages using read_document and counts keyword hits for patent examination terminology. It returns a hint field ("likely", "unclear") based on the density of relevant terms, allowing administrators to filter books before resource-intensive processing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →