# How to Ingest Historical Cases into the Response Assistance Database: A Complete Guide

> Learn to ingest historical cases into the response assistance database. This guide covers parsing Markdown, generating vector embeddings, and persisting to SQLite for your patent disclosure skill.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-01

---

**To ingest historical cases into the response assistance database, parse case Markdown files into `Case` objects, generate vector embeddings, and persist them to SQLite with [`tools/oa/ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/ingest_case.py), then rebuild the search index.**

The **patent-disclosure-skill** repository by `handsomestWei` implements a modular "Open-Assistance" (OA) framework that powers patent response assistance through vector-based semantic search. Ingesting historical cases follows a pipeline architecture spanning parsing, embedding generation, and indexed storage. This article walks through each step using the actual source implementation.

## Understanding the Ingestion Pipeline Architecture

The OA framework separates concerns across specialized modules. At a high level, historical case ingestion involves five distinct operations coordinated through well-defined interfaces.

The pipeline flow:

1. **Case Definition** – Markdown files describe patents, claims, and analysis
2. **Parsing & Normalization** – Extract metadata and structured content
3. **Vectorization** – Generate dense embeddings for semantic search
4. **Persistence** – Store cases and vectors in SQLite
5. **Index Refresh** – Rebuild approximate nearest neighbor indexes for fast retrieval

This architecture lives under the `tools/oa/` directory, with each pipeline stage implemented as a focused Python module.

## Step 1: Prepare Historical Case Files in the Vault

Historical cases reside as **Markdown files** in the OA vault. The vault structure is defined in [`tools/oa/vault_layout.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/vault_layout.py).

Each case file should include:
- YAML frontmatter with metadata (title, tags, publication date)
- Markdown body with patent claims and analysis

Example file structure at [`cases/US1234567_A1.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cases/US1234567_A1.md):

```markdown
---
title: "Novel Wireless Charging Coil"
tags: ["wireless-power", "induction", "consumer-electronics"]
publication_date: "2023-08-15"
application_number: "US16/123,456"
---

## Claims

1. A wireless charging apparatus comprising...

## Analysis

Prior art search reveals similar coils in US8,123,456, but the present disclosure distinguishes through...

```

Place all case files in a dedicated folder within the vault before invoking ingestion.

## Step 2: Parse Cases with [`ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/ingest_case.py)

The [`tools/oa/ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/ingest_case.py) module provides the core ingestion routine. It transforms Markdown files into structured `Case` objects using [`tools/oa/case_md.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/case_md.py).

Key operations in `ingest_case()`:

```python
def ingest_case(md_path: Path):
    case = Case.from_markdown(md_path)          # parse Markdown → Case object

    vector = embed_text(case.full_text)         # generate embedding

    Store.add_case(case, vector)                # persist case & vector

```

The `Case.from_markdown()` factory method in [`tools/oa/case_md.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/case_md.py) handles:
- YAML frontmatter extraction
- Body text normalization
- Metadata validation

### Single Case Ingestion Example

```python
from tools.oa.ingest_case import ingest_case
from pathlib import Path

case_path = Path("cases/US1234567_A1.md")
ingest_case(case_path)

```

This one-line call executes the full pipeline: parsing, embedding, and storage.

## Step 3: Generate Vector Embeddings

The [`tools/oa/embed.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/embed.py) module wraps the configured embedding service. It converts case text into dense vectors using models specified in [`tools/oa/config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/config.py).

Supported embedding backends include:
- **OpenAI** (text-embedding-ada-002, text-embedding-3-small)
- **Cohere** embed models
- Local sentence-transformers (configurable)

The `embed_text()` function:
- Chunks long documents if needed
- Handles rate limiting and retries
- Returns normalized vectors for consistent similarity scoring

Vectors are dimension-aligned with the search index (typically 768 or 1536 dimensions depending on the model).

## Step 4: Persist Cases to SQLite Storage

The [`tools/oa/store.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/store.py) module implements an **SQLite-backed persistence layer** with two core tables:
- `cases` – structured case metadata and content
- `vectors` – embedding vectors with foreign key references

The `Store.add_case()` method:
- Inserts case record with full text and metadata
- Stores associated vector with case ID
- Maintains transaction integrity

This dual-table design enables:
- Fast metadata filtering (`SELECT * FROM cases WHERE tags LIKE '%wireless%'`)
- Efficient vector similarity search via indexed ANN structures

## Step 5: Bulk Ingestion with [`ingest_playbook.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/ingest_playbook.py)

For production deployments, [`tools/oa/ingest_playbook.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/ingest_playbook.py) provides batch processing capabilities.

### Bulk Folder Ingestion Example

```python
from tools.oa.ingest_playbook import ingest_folder

ingest_folder("cases/")

```

This helper automatically:
- Discovers all `.md` files in the target directory
- Loops through `ingest_case()` for each file
- Handles errors gracefully with per-file logging
- Reports ingestion statistics on completion

The playbook is the recommended entry point for initial database population and periodic batch updates.

## Step 6: Rebuild the Vector Search Index

After ingestion, new cases require index rebuild for searchability. Run [`tools/oa/rebuild_vectors.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/rebuild_vectors.py) to synchronize the approximate nearest neighbor index.

```bash
python -m tools.oa.rebuild_vectors

```

This script:
- Reads all vectors from [`tools/oa/store.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/store.py)
- Reconstructs the FAISS or Annoy index (backend configurable in [`tools/oa/config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/config.py))
- Persists index to disk for fast loading at assistant startup

**Critical:** Skip this step and new cases exist in storage but remain invisible to similarity queries.

## Key Source Files Reference

| File | Purpose |
|------|---------|
| [`tools/oa/ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/ingest_case.py) | Core ingestion routine: parse, embed, store single case |
| [`tools/oa/case_md.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/case_md.py) | `Case` data model with `from_markdown` factory |
| [`tools/oa/embed.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/embed.py) | Embedding service wrapper with model abstraction |
| [`tools/oa/store.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/store.py) | SQLite persistence: `add_case()`, `get_case()`, vector storage |
| [`tools/oa/rebuild_vectors.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/rebuild_vectors.py) | Index reconstruction for similarity search |
| [`tools/oa/ingest_playbook.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/ingest_playbook.py) | Bulk directory ingestion orchestration |
| [`tools/oa/vault_layout.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/vault_layout.py) | Vault directory structure conventions |
| [`tools/oa/config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/config.py) | Embedding model and backend configuration |

## Production Ingestion Workflow

A complete historical case migration follows this sequence:

```bash

# 1. Verify vault structure

python -m tools.oa.vault_layout --validate cases/

# 2. Bulk ingest all historical cases

python -m tools.oa.ingest_playbook --source cases/ --verbose

# 3. Rebuild search index

python -m tools.oa.rebuild_vectors

# 4. Verify ingestion count

python -c "from tools.oa.store import Store; print(Store.case_count())"

```

## Summary

- **Place case Markdown files** in the OA vault following [`tools/oa/vault_layout.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/vault_layout.py) conventions
- **Use `ingest_case()`** from [`tools/oa/ingest_case.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/ingest_case.py) for single-case ingestion with automatic parsing, embedding, and storage
- **Leverage [`ingest_playbook.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/ingest_playbook.py)** for batch processing entire directories
- **Always run [`rebuild_vectors.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/rebuild_vectors.py)** after ingestion to make new cases searchable
- **Configure embedding models** in [`tools/oa/config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/config.py) to match your semantic search requirements

## Frequently Asked Questions

### What Markdown format must historical case files follow?

Historical case files require YAML frontmatter with `title`, `tags`, and `publication_date` fields, followed by standard Markdown body text containing claims and analysis. The `Case.from_markdown()` method in [`tools/oa/case_md.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/case_md.py) parses this structure and validates required fields.

### Can I use a different embedding model without changing code?

Yes. Model selection is configured in [`tools/oa/config.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/config.py). Change the embedding provider and model name there; [`tools/oa/embed.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/embed.py) abstracts the implementation so ingestion code remains unchanged.

### Why must I rebuild the vector index after ingestion?

The FAISS or Annoy index used for similarity search is a static data structure written to disk. New vectors in SQLite aren't automatically added to the loaded index. [`tools/oa/rebuild_vectors.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/rebuild_vectors.py) reconstructs the index from all stored vectors, ensuring complete search coverage.

### How do I verify cases were ingested successfully?

Query `Store.case_count()` from [`tools/oa/store.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/tools/oa/store.py) for a total count, or retrieve specific cases with `Store.get_case(case_id)`. For search functionality verification, run a similarity query against a known case title and confirm the expected result ranks highest.