How to Ingest Historical Cases into the Response Assistance Database: A Complete Guide
To ingest historical cases into the response assistance database, parse case Markdown files into Case objects, generate vector embeddings, and persist them to SQLite with tools/oa/ingest_case.py, then rebuild the search index.
The patent-disclosure-skill repository by handsomestWei implements a modular "Open-Assistance" (OA) framework that powers patent response assistance through vector-based semantic search. Ingesting historical cases follows a pipeline architecture spanning parsing, embedding generation, and indexed storage. This article walks through each step using the actual source implementation.
Understanding the Ingestion Pipeline Architecture
The OA framework separates concerns across specialized modules. At a high level, historical case ingestion involves five distinct operations coordinated through well-defined interfaces.
The pipeline flow:
- Case Definition – Markdown files describe patents, claims, and analysis
- Parsing & Normalization – Extract metadata and structured content
- Vectorization – Generate dense embeddings for semantic search
- Persistence – Store cases and vectors in SQLite
- Index Refresh – Rebuild approximate nearest neighbor indexes for fast retrieval
This architecture lives under the tools/oa/ directory, with each pipeline stage implemented as a focused Python module.
Step 1: Prepare Historical Case Files in the Vault
Historical cases reside as Markdown files in the OA vault. The vault structure is defined in tools/oa/vault_layout.py.
Each case file should include:
- YAML frontmatter with metadata (title, tags, publication date)
- Markdown body with patent claims and analysis
Example file structure at cases/US1234567_A1.md:
---
title: "Novel Wireless Charging Coil"
tags: ["wireless-power", "induction", "consumer-electronics"]
publication_date: "2023-08-15"
application_number: "US16/123,456"
---
## Claims
1. A wireless charging apparatus comprising...
## Analysis
Prior art search reveals similar coils in US8,123,456, but the present disclosure distinguishes through...
Place all case files in a dedicated folder within the vault before invoking ingestion.
Step 2: Parse Cases with ingest_case.py
The tools/oa/ingest_case.py module provides the core ingestion routine. It transforms Markdown files into structured Case objects using tools/oa/case_md.py.
Key operations in ingest_case():
def ingest_case(md_path: Path):
case = Case.from_markdown(md_path) # parse Markdown → Case object
vector = embed_text(case.full_text) # generate embedding
Store.add_case(case, vector) # persist case & vector
The Case.from_markdown() factory method in tools/oa/case_md.py handles:
- YAML frontmatter extraction
- Body text normalization
- Metadata validation
Single Case Ingestion Example
from tools.oa.ingest_case import ingest_case
from pathlib import Path
case_path = Path("cases/US1234567_A1.md")
ingest_case(case_path)
This one-line call executes the full pipeline: parsing, embedding, and storage.
Step 3: Generate Vector Embeddings
The tools/oa/embed.py module wraps the configured embedding service. It converts case text into dense vectors using models specified in tools/oa/config.py.
Supported embedding backends include:
- OpenAI (text-embedding-ada-002, text-embedding-3-small)
- Cohere embed models
- Local sentence-transformers (configurable)
The embed_text() function:
- Chunks long documents if needed
- Handles rate limiting and retries
- Returns normalized vectors for consistent similarity scoring
Vectors are dimension-aligned with the search index (typically 768 or 1536 dimensions depending on the model).
Step 4: Persist Cases to SQLite Storage
The tools/oa/store.py module implements an SQLite-backed persistence layer with two core tables:
cases– structured case metadata and contentvectors– embedding vectors with foreign key references
The Store.add_case() method:
- Inserts case record with full text and metadata
- Stores associated vector with case ID
- Maintains transaction integrity
This dual-table design enables:
- Fast metadata filtering (
SELECT * FROM cases WHERE tags LIKE '%wireless%') - Efficient vector similarity search via indexed ANN structures
Step 5: Bulk Ingestion with ingest_playbook.py
For production deployments, tools/oa/ingest_playbook.py provides batch processing capabilities.
Bulk Folder Ingestion Example
from tools.oa.ingest_playbook import ingest_folder
ingest_folder("cases/")
This helper automatically:
- Discovers all
.mdfiles in the target directory - Loops through
ingest_case()for each file - Handles errors gracefully with per-file logging
- Reports ingestion statistics on completion
The playbook is the recommended entry point for initial database population and periodic batch updates.
Step 6: Rebuild the Vector Search Index
After ingestion, new cases require index rebuild for searchability. Run tools/oa/rebuild_vectors.py to synchronize the approximate nearest neighbor index.
python -m tools.oa.rebuild_vectors
This script:
- Reads all vectors from
tools/oa/store.py - Reconstructs the FAISS or Annoy index (backend configurable in
tools/oa/config.py) - Persists index to disk for fast loading at assistant startup
Critical: Skip this step and new cases exist in storage but remain invisible to similarity queries.
Key Source Files Reference
| File | Purpose |
|---|---|
tools/oa/ingest_case.py |
Core ingestion routine: parse, embed, store single case |
tools/oa/case_md.py |
Case data model with from_markdown factory |
tools/oa/embed.py |
Embedding service wrapper with model abstraction |
tools/oa/store.py |
SQLite persistence: add_case(), get_case(), vector storage |
tools/oa/rebuild_vectors.py |
Index reconstruction for similarity search |
tools/oa/ingest_playbook.py |
Bulk directory ingestion orchestration |
tools/oa/vault_layout.py |
Vault directory structure conventions |
tools/oa/config.py |
Embedding model and backend configuration |
Production Ingestion Workflow
A complete historical case migration follows this sequence:
# 1. Verify vault structure
python -m tools.oa.vault_layout --validate cases/
# 2. Bulk ingest all historical cases
python -m tools.oa.ingest_playbook --source cases/ --verbose
# 3. Rebuild search index
python -m tools.oa.rebuild_vectors
# 4. Verify ingestion count
python -c "from tools.oa.store import Store; print(Store.case_count())"
Summary
- Place case Markdown files in the OA vault following
tools/oa/vault_layout.pyconventions - Use
ingest_case()fromtools/oa/ingest_case.pyfor single-case ingestion with automatic parsing, embedding, and storage - Leverage
ingest_playbook.pyfor batch processing entire directories - Always run
rebuild_vectors.pyafter ingestion to make new cases searchable - Configure embedding models in
tools/oa/config.pyto match your semantic search requirements
Frequently Asked Questions
What Markdown format must historical case files follow?
Historical case files require YAML frontmatter with title, tags, and publication_date fields, followed by standard Markdown body text containing claims and analysis. The Case.from_markdown() method in tools/oa/case_md.py parses this structure and validates required fields.
Can I use a different embedding model without changing code?
Yes. Model selection is configured in tools/oa/config.py. Change the embedding provider and model name there; tools/oa/embed.py abstracts the implementation so ingestion code remains unchanged.
Why must I rebuild the vector index after ingestion?
The FAISS or Annoy index used for similarity search is a static data structure written to disk. New vectors in SQLite aren't automatically added to the loaded index. tools/oa/rebuild_vectors.py reconstructs the index from all stored vectors, ensuring complete search coverage.
How do I verify cases were ingested successfully?
Query Store.case_count() from tools/oa/store.py for a total count, or retrieve specific cases with Store.get_case(case_id). For search functionality verification, run a similarity query against a known case title and confirm the expected result ranks highest.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →