How the Office Action Response Mode Builds a RAG Case Library with Optional Vector Embeddings

The office action response workflow in handsomestWei/patent-disclosure-skill constructs a retrieval-augmented generation (RAG) case library by scanning Markdown case files, building a keyword index, and optionally generating vector embeddings for semantic search.

This self-contained RAG system enables patent practitioners to retrieve relevant precedent cases when drafting responses to examiner office actions. The implementation resides entirely within the tools/oa/ directory and operates without external database dependencies unless vector embeddings are explicitly enabled.

Case Ingestion: From Markdown Files to Structured Records

The foundation of the RAG library is the case ingestion pipeline. The tools/oa/case_md.py module handles parsing of standardized Markdown case files, extracting both metadata and substantive content.

The parse_case_markdown function reads each case file and splits it into front-matter metadata and body text:


# tools/oa/case_md.py – parse a case markdown file

def parse_case_markdown(text: str) -> tuple[dict, str]:
    meta, body = yaml.safe_load_front_matter(text), extract_body(text)
    # meta contains slug, title, tags, source_path …

    return meta, body

The companion dump_case_markdown function serializes records back to the canonical Markdown format. These utilities enable bidirectional transformation between human-editable case files and program-readable data structures.

Each case file typically contains:

  • slug: Unique identifier for the case
  • title: Human-readable case title
  • tags: Categorical labels (e.g., "缺陷", "法条", "实用新型")
  • source_path: Reference to original PDF or source document
  • body: The distilled legal reasoning and response strategy

Index Construction: Building the Searchable Case Library

The tools/oa/playbook.py module orchestrates case discovery and index generation. The list_playbook_records function walks the case directory structure and builds the in-memory representation of the library:


# tools/oa/playbook.py – collect all case records

def list_playbook_records(oa_root: Path) -> list[dict]:
    root = playbooks_root(oa_root)          # = oa/playbooks

    records = []
    for d in sorted(root.iterdir()):
        if d.is_dir():
            index = d / PLAYBOOK_INDEX
            meta = {}
            if index.is_file():
                meta, _ = parse_case_markdown(index.read_text())
            records.append({
                "slug": str(meta.get("slug") or d.name),
                "title": str(meta.get("title") or d.name),
                "tags": meta.get("tags", []),
                "source_path": str(meta.get("source_path") or ""),
            })
    return records

This function populates the RECORDS list that serves as the primary data structure for all subsequent retrieval operations. The index file _playbook.md (referenced by the PLAYBOOK_INDEX constant) acts as the canonical metadata store for each case directory.

The ingest_distilled_skill function (invoked when processing new cases) writes these index files and ensures consistency between the source Markdown and the searchable library representation.

The RAG architecture supports an optional vector embedding layer for semantic similarity search. This capability is controlled by the use_vector configuration flag in tools/oa/config.yaml and implemented in tools/oa/vector_index.py.

When enabled, the embed_text function transforms case content into dense vector representations:


# tools/oa/vector_index.py – optional embedding step

def embed_text(text: str) -> List[float]:
    import openai
    resp = openai.Embedding.create(
        model="text-embedding-ada-002",
        input=text
    )
    return resp["data"][0]["embedding"]

The embedding process concatenates the case title and body fields to capture both topical and substantive semantic information. The resulting 1536-dimensional vectors are stored in a local JSON file (vector_store.json) keyed by case slug, eliminating the need for external vector databases.

Vector generation occurs lazily—only when cases are added or the embedding flag is toggled—keeping the default installation lightweight and privacy-preserving.

Retrieval Architecture: Hybrid Search with Graceful Degradation

The query-time retrieval logic in tools/oa/search.py implements a hybrid strategy that prioritizes vector search when available and falls back to keyword matching otherwise:


# tools/oa/search.py – combined retrieval

def retrieve_cases(query: str, use_vector: bool = False):
    if use_vector and VECTOR_STORE:
        q_vec = embed_text(query)
        hits = nearest_vectors(q_vec, VECTOR_STORE, k=5)
        slugs = [hit.slug for hit in hits]
    else:
        slugs = [r["slug"] for r in RECORDS if any(t in query for t in r["tags"])]
    return [load_case(slug) for slug in slugs]

The nearest_vectors helper performs a linear scan against the cached vector store—a design choice appropriate for the modest case library sizes typical of specialized patent domains. For larger collections, the modular structure permits drop-in replacement with approximate nearest neighbor libraries like FAISS or Annoy.

The fallback tag-based search uses simple substring matching against the tags field, providing deterministic retrieval without external API calls.

Configuration and Activation

Vector embedding support is opt-in by design. Users enable the feature through the OA configuration file:


# tools/oa/config.yaml

use_vector: true
embedding_model: text-embedding-ada-002
vector_store_path: vector_store.json

When use_vector is false or absent, the system operates in pure keyword mode with no OpenAI API calls and no local vector storage. This architecture ensures that:

  • Default installations remain offline-capable and zero-cost
  • Sensitive patent content never leaves the local environment unless explicitly configured
  • Vector capabilities can be toggled without structural changes to the case library

Summary

  • The office action response mode builds its RAG library by parsing Markdown case files via tools/oa/case_md.py, extracting structured metadata and body content.
  • Index construction happens through tools/oa/playbook.py, which discovers cases and maintains the RECORDS data structure for fast lookup.
  • Vector embeddings are strictly optional—controlled by use_vector in config.yaml and implemented in tools/oa/vector_index.py using OpenAI's embedding API with local JSON persistence.
  • Retrieval operates in hybrid mode: vector similarity when enabled, tag-based keyword matching as the universal fallback.

Frequently Asked Questions

What file formats does the case library support?

The system ingests Markdown files with YAML front-matter as the canonical case format. Each case resides in its own directory under oa/cases/ or oa/playbooks/, containing a Markdown index file plus optional referenced assets like source PDFs. The parse_case_markdown function in tools/oa/case_md.py standardizes extraction of metadata fields regardless of minor formatting variations.

How does the system handle cases without vector embeddings?

Cases without embeddings participate fully in tag-based retrieval. The retrieve_cases function in tools/oa/search.py inspects the VECTOR_STORE global and automatically degrades to substring matching against the tags field when vectors are unavailable. All cases remain discoverable; only the ranking mechanism changes.

Can I use a different embedding model or provider?

The tools/oa/vector_index.py module encapsulates all embedding operations behind the embed_text function. Replacing the OpenAI implementation requires modifying only this function—swap the API call for your preferred provider (Cohere, local Hugging Face models, etc.) and adjust vector dimensionality accordingly. The rest of the retrieval pipeline remains unchanged.

Where is the vector store physically stored?

Vector embeddings persist to a local JSON file specified by vector_store_path in config.yaml (default: vector_store.json in the tools/oa/ directory). This file contains a dictionary mapping case slugs to float arrays. No external databases, cloud services, or Docker containers are required.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →