# How CNIPA Search Results Are Formatted and Saved in patent-disclosure-skill

> Discover how CNIPA search results are transformed into structured EpubSearchHit objects, serialized to JSON, and saved as Markdown reports in the patent-disclosure-skill project.

- Repository: [handsomestWei/patent-disclosure-skill](https://github.com/handsomestWei/patent-disclosure-skill)
- Tags: how-to-guide
- Published: 2026-09-06

---

**CNIPA search results are extracted from raw HTML into structured `EpubSearchHit` objects, serialized to a JSON payload, and rendered as a Markdown report saved to the `outputs/patent-search` directory by default.**

The `handsomestWei/patent-disclosure-skill` repository automates queries against the China National Intellectual Property Administration (CNIPA) database at `http://epub.cnipa.gov.cn`. Understanding how CNIPA search results are structured and persisted enables seamless integration into patent research and disclosure workflows.

## Data Structure and Parsing

In [`skills/patent-search/tools/cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_parse.py), the `parse_search_result_html` function transforms raw HTML from the CNIPA advanced query page into a list of **`EpubSearchHit`** dataclass instances. Each hit contains the following optional fields extracted from the search result rows:

- **title**: Patent name or title text
- **pub_number**: Publication/announcement number (e.g., CN123456789A)
- **application_number**: Normalized Chinese application number
- **applicant**: Patent right holder or applicant name
- **inventors**: List of inventor names
- **filing_date**: Application filing date
- **publication_date**: Publication or announcement date
- **link**: Absolute URL to the detailed patent page
- **abstract**: Full abstract text when available
- **ipc_codes**: List of IPC classification codes
- **loc_codes**: List of Locarno (LOC) classification codes
- **raw_html**: Truncated HTML snippet retained for debugging purposes

## JSON Payload and Markdown Report Generation

After parsing, [`skills/patent-search/tools/cnipa_search.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_search.py) assembles the hits into a plain JSON payload. This payload is passed to **`render_search_report`** in [`skills/patent-search/tools/emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/emit_search_report.py), which produces a structured Markdown document containing:

1. A header displaying the query timestamp and source URL
2. A parameter table listing the search criteria used
3. Pagination metadata (pages scanned, total pages, and page size)
4. Numbered sections for each hit, formatted as "`1. [Title](URL)`" followed by bullet points showing all extracted fields, classification codes, and the abstract text

The **`write_search_report`** function handles the final file I/O operation, writing the rendered Markdown to disk.

## Default Output Location and Configuration

By default, `write_search_report` saves files to the repository's **`outputs/patent-search`** directory using the timestamped naming convention:

```text
SEARCH-YYYYMMDD-HHMMSS.md

```

You can override this location by setting the **`PATENT_SEARCH_OUTPUT_DIR`** environment variable before execution. When running the CLI tool, the system outputs two environment variables for downstream automation:

- `EPUB_SEARCH_MD`: Absolute path to the generated Markdown file
- `EPUB_SEARCH_JSON`: JSON payload printed to stdout

## Practical Code Examples

Here is how to parse saved HTML and generate a report programmatically:

```python
from pathlib import Path
from cnipa_parse import parse_search_result_html
from emit_search_report import write_search_report

# Parse HTML into structured hits

html = Path("example.html").read_text(encoding="utf-8")
hits = parse_search_result_html(html)  # Returns List[EpubSearchHit]

# Prepare payload

payload = {
    "source": "http://epub.cnipa.gov.cn/Advanced",
    "query": {"inventor": "张三"},
    "hits": [hit.__dict__ for hit in hits],
    # Additional metadata added by cnipa_search.py

}

# Generate and save report

report_path = write_search_report(payload)
print(f"Report saved to {report_path}")

```

To run the complete pipeline from the command line:

```bash
python skills/patent-search/tools/cnipa_search.py \
    --inventor "张三" \
    --applicant "北京大学" \
    --max-pages 2

```

## Summary

- **CNIPA search results** originate as raw HTML from `epub.cnipa.gov.cn` and are parsed into `EpubSearchHit` objects by [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py).
- **Data extraction** captures 12 distinct fields including publication numbers, dates, applicants, inventors, and classification codes.
- **Report generation** converts structured data to JSON and renders it as Markdown via [`emit_search_report.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/emit_search_report.py).
- **Default save location** is [`outputs/patent-search/SEARCH-YYYYMMDD-HHMMSS.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/outputs/patent-search/SEARCH-YYYYMMDD-HHMMSS.md), configurable via the `PATENT_SEARCH_OUTPUT_DIR` environment variable.
- **CLI integration** exposes file paths through `EPUB_SEARCH_MD` and `EPUB_SEARCH_JSON` environment variables for automated workflow chaining.

## Frequently Asked Questions

### What data fields are extracted from CNIPA search results?

According to the [`cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/cnipa_parse.py) implementation, the pipeline extracts title, publication number, application number, applicant, inventors, filing date, publication date, link, abstract, IPC codes, LOC codes, and raw HTML snippets. All fields are optional and depend on availability in the source HTML returned by the CNIPA server.

### How can I change where the CNIPA search reports are saved?

Set the `PATENT_SEARCH_OUTPUT_DIR` environment variable to your desired directory path before running the search. If this variable is not set, `write_search_report` defaults to the `outputs/patent-search` folder relative to the repository root, using the [`SEARCH-YYYYMMDD-HHMMSS.md`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/SEARCH-YYYYMMDD-HHMMSS.md) filename pattern.

### What is the difference between the JSON and Markdown outputs?

The JSON payload (referenced by `EPUB_SEARCH_JSON` and printed to stdout) contains the raw structured data with all `EpubSearchHit` fields serialized as dictionaries. The Markdown report (referenced by `EPUB_SEARCH_MD`) is a human-readable document generated by `render_search_report` that formats this data into tables, bullet points, and sections for easy review and sharing.

### Which file handles the conversion of HTML to structured data?

The [`skills/patent-search/tools/cnipa_parse.py`](https://github.com/handsomestWei/patent-disclosure-skill/blob/main/skills/patent-search/tools/cnipa_parse.py) module contains the `parse_search_result_html` function responsible for scraping the CNIPA HTML and instantiating `EpubSearchHit` objects with the extracted patent metadata, including pagination information for multi-page result sets.