How CNIPA Search Results Are Formatted and Saved in patent-disclosure-skill

CNIPA search results are extracted from raw HTML into structured EpubSearchHit objects, serialized to a JSON payload, and rendered as a Markdown report saved to the outputs/patent-search directory by default.

The handsomestWei/patent-disclosure-skill repository automates queries against the China National Intellectual Property Administration (CNIPA) database at http://epub.cnipa.gov.cn. Understanding how CNIPA search results are structured and persisted enables seamless integration into patent research and disclosure workflows.

Data Structure and Parsing

In skills/patent-search/tools/cnipa_parse.py, the parse_search_result_html function transforms raw HTML from the CNIPA advanced query page into a list of EpubSearchHit dataclass instances. Each hit contains the following optional fields extracted from the search result rows:

  • title: Patent name or title text
  • pub_number: Publication/announcement number (e.g., CN123456789A)
  • application_number: Normalized Chinese application number
  • applicant: Patent right holder or applicant name
  • inventors: List of inventor names
  • filing_date: Application filing date
  • publication_date: Publication or announcement date
  • link: Absolute URL to the detailed patent page
  • abstract: Full abstract text when available
  • ipc_codes: List of IPC classification codes
  • loc_codes: List of Locarno (LOC) classification codes
  • raw_html: Truncated HTML snippet retained for debugging purposes

JSON Payload and Markdown Report Generation

After parsing, skills/patent-search/tools/cnipa_search.py assembles the hits into a plain JSON payload. This payload is passed to render_search_report in skills/patent-search/tools/emit_search_report.py, which produces a structured Markdown document containing:

  1. A header displaying the query timestamp and source URL
  2. A parameter table listing the search criteria used
  3. Pagination metadata (pages scanned, total pages, and page size)
  4. Numbered sections for each hit, formatted as "1. [Title](URL)" followed by bullet points showing all extracted fields, classification codes, and the abstract text

The write_search_report function handles the final file I/O operation, writing the rendered Markdown to disk.

Default Output Location and Configuration

By default, write_search_report saves files to the repository's outputs/patent-search directory using the timestamped naming convention:

SEARCH-YYYYMMDD-HHMMSS.md

You can override this location by setting the PATENT_SEARCH_OUTPUT_DIR environment variable before execution. When running the CLI tool, the system outputs two environment variables for downstream automation:

  • EPUB_SEARCH_MD: Absolute path to the generated Markdown file
  • EPUB_SEARCH_JSON: JSON payload printed to stdout

Practical Code Examples

Here is how to parse saved HTML and generate a report programmatically:

from pathlib import Path
from cnipa_parse import parse_search_result_html
from emit_search_report import write_search_report

# Parse HTML into structured hits

html = Path("example.html").read_text(encoding="utf-8")
hits = parse_search_result_html(html)  # Returns List[EpubSearchHit]

# Prepare payload

payload = {
    "source": "http://epub.cnipa.gov.cn/Advanced",
    "query": {"inventor": "张三"},
    "hits": [hit.__dict__ for hit in hits],
    # Additional metadata added by cnipa_search.py

}

# Generate and save report

report_path = write_search_report(payload)
print(f"Report saved to {report_path}")

To run the complete pipeline from the command line:

python skills/patent-search/tools/cnipa_search.py \
    --inventor "张三" \
    --applicant "北京大学" \
    --max-pages 2

Summary

  • CNIPA search results originate as raw HTML from epub.cnipa.gov.cn and are parsed into EpubSearchHit objects by cnipa_parse.py.
  • Data extraction captures 12 distinct fields including publication numbers, dates, applicants, inventors, and classification codes.
  • Report generation converts structured data to JSON and renders it as Markdown via emit_search_report.py.
  • Default save location is outputs/patent-search/SEARCH-YYYYMMDD-HHMMSS.md, configurable via the PATENT_SEARCH_OUTPUT_DIR environment variable.
  • CLI integration exposes file paths through EPUB_SEARCH_MD and EPUB_SEARCH_JSON environment variables for automated workflow chaining.

Frequently Asked Questions

What data fields are extracted from CNIPA search results?

According to the cnipa_parse.py implementation, the pipeline extracts title, publication number, application number, applicant, inventors, filing date, publication date, link, abstract, IPC codes, LOC codes, and raw HTML snippets. All fields are optional and depend on availability in the source HTML returned by the CNIPA server.

How can I change where the CNIPA search reports are saved?

Set the PATENT_SEARCH_OUTPUT_DIR environment variable to your desired directory path before running the search. If this variable is not set, write_search_report defaults to the outputs/patent-search folder relative to the repository root, using the SEARCH-YYYYMMDD-HHMMSS.md filename pattern.

What is the difference between the JSON and Markdown outputs?

The JSON payload (referenced by EPUB_SEARCH_JSON and printed to stdout) contains the raw structured data with all EpubSearchHit fields serialized as dictionaries. The Markdown report (referenced by EPUB_SEARCH_MD) is a human-readable document generated by render_search_report that formats this data into tables, bullet points, and sections for easy review and sharing.

Which file handles the conversion of HTML to structured data?

The skills/patent-search/tools/cnipa_parse.py module contains the parse_search_result_html function responsible for scraping the CNIPA HTML and instantiating EpubSearchHit objects with the extracted patent metadata, including pagination information for multi-page result sets.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →