How the CAD/STEP Parsing Pipeline Extracts Multi-View Projections Without Hindering Disclosure Drafting

The CAD/STEP parsing pipeline generates orthogonal view images entirely in memory and returns them as base-64 PNG data, enabling seamless embedding into patent drafts without file-system side effects or geometric modifications.

The handsomestWei/patent-disclosure-skill repository implements a self-contained CAD processing system that transforms STEP files into multi-view projections suitable for patent documentation. This article examines how the pipeline balances geometric fidelity with drafting workflow requirements.

STEP File Ingestion and Model Wrapping

The entry point for CAD processing resides in tools/shared/cad_scan.py. This module leverages python-occ, the Python bindings for OpenCascade Technology (OCC), to read STEP files without modifying their contents.

The load_step() function performs three operations:

  • Opens the STEP file through OCC's STEPControl_Reader
  • Transfers the root shape into memory as a TopoDS_Shape object
  • Wraps the shape in a lightweight CADModel dataclass for downstream consumption

# tools/shared/cad_scan.py (conceptual)

from OCC.Core.STEPControl import STEPControl_Reader
from dataclasses import dataclass

@dataclass
class CADModel:
    shape: TopoDS_Shape
    filename: str

def load_step(path: str) -> CADModel:
    reader = STEPControl_Reader()
    status = reader.ReadFile(path)
    if status != IFSelect_RetDone:
        raise ValueError(f"Failed to read STEP: {path}")
    reader.TransferRoots()
    shape = reader.OneShape()
    return CADModel(shape=shape, filename=path)

Because load_step() operates read-only, the source STEP file remains untouched throughout the disclosure drafting process.

Multi-View Projection Generation

The core projection logic lives in tools/shared/step_to_views.py. This module defines canonical view orientations and renders each projection to an off-screen buffer.

View Matrix Configuration


# tools/shared/step_to_views.py (conceptual)

VIEWS = {
    "front":  (1, 0, 0, 0, 0, -1, 0, 1, 0),   # Looking along +Y

    "top":    (1, 0, 0, 0, 0, 1, 0, 0, -1),   # Looking along -Z

    "right":  (0, 0, 1, 0, 1, 0, -1, 0, 0),  # Looking along +X

    "iso":    (0.707, 0.5, 0.5, -0.707, 0.5, 0.5, 0, -0.707, 0.707)
}

Each 9-tuple represents a 3×3 rotation matrix that orients the virtual camera relative to the model's coordinate system.

In-Memory Rendering Pipeline

The render_views() function processes the CADModel through a pipeline that avoids persistent storage:


# tools/shared/step_to_views.py (conceptual)

import base64
from io import BytesIO
from OCC.Core.V3d import V3d_Viewer
from OCC.Core.Graphic3d import Graphic3d_RenderingParams

def render_views(cad_model: CADModel, size: tuple = (512, 512)) -> dict:
    """
    Generate base-64 encoded PNGs for each standard projection.
    """
    images = {}
    
    for view_name, view_matrix in VIEWS.items():
        # Configure off-screen viewer

        viewer = V3d_Viewer(Graphic3d_RenderingParams())
        view = viewer.CreateView()
        view.SetSize(*size)
        view.SetBgGradientColors(
            Quantity_Color(1, 1, 1, Quantity_TOC_RGB),
            Quantity_Color(1, 1, 1, Quantity_TOC_RGB),
            0, True
        )
        
        # Apply projection matrix

        camera = view.Camera()
        camera.SetDirection(view_matrix[3], view_matrix[4], view_matrix[5])
        camera.SetUp(view_matrix[6], view_matrix[7], view_matrix[8])
        
        # Render to raster

        view.Display(cad_model.shape)
        view.FitAll()
        
        # Capture to PNG bytes in memory

        buffer = BytesIO()
        view.Dump(buffer, size[0], size[1])
        png_bytes = buffer.getvalue()
        
        # Encode as data URL for immediate embedding

        b64 = base64.b64encode(png_bytes).decode('ascii')
        images[view_name] = f"data:image/png;base64,{b64}"
    
    return images

Key design decisions that preserve drafting workflow integrity:

  • No temporary files: The BytesIO buffer holds raster data exclusively in RAM
  • Fixed resolution: Default 512×512 pixels balances detail against payload size
  • Pure function output: The returned dictionary contains only strings, making it serializable and cache-friendly

Integration with Disclosure Drafting

The tools/shared/run_step_to_views.py module serves as the CLI bridge between CAD processing and document generation. It emits a JSON structure that downstream tools consume directly.


# tools/shared/run_step_to_views.py (conceptual)

import json
import sys
from cad_scan import load_step
from step_to_views import render_views

def main():
    step_path = sys.argv[1]
    model = load_step(step_path)
    views = render_views(model)
    
    payload = {
        "source_file": model.filename,
        "views": views,
        "count": len(views)
    }
    sys.stdout.write(json.dumps(payload, indent=2))

if __name__ == "__main__":
    main()

The JSON output integrates with tools/shared/md_to_docx.py (the markdown-to-docx converter) through standard input/output streams. This architecture ensures that:

  • The drafting process never blocks on CAD operations
  • Image data flows as opaque strings, preventing accidental mutation
  • The original STEP file path is preserved for citation purposes but not accessed during document assembly

Why the Pipeline Preserves Drafting Workflow Efficiency

Risk Factor Mitigation Strategy Implementation Location
File-system contamination In-memory rendering with BytesIO step_to_views.py:render_views()
Geometric drift Read-only STEP access; no shape modification cad_scan.py:load_step()
Format incompatibility Base-64 PNG output universally supported step_to_views.py data URL encoding
Performance bottlenecks Lightweight 512×512 renders; synchronous execution render_views() default parameters
Blocking I/O JSON streaming to stdout run_step_to_views.py:main()

Practical Usage Example

Below is a complete workflow demonstrating CAD integration into a disclosure draft:

import json
import subprocess

def embed_cad_views(step_path: str, draft_md: str) -> str:
    """
    Inject multi-view projections into a markdown disclosure draft.
    """
    # Execute pipeline as subprocess to isolate OCC dependencies

    result = subprocess.run(
        ["python", "tools/shared/run_step_to_views.py", step_path],
        capture_output=True, text=True, check=True
    )
    views_data = json.loads(result.stdout)
    
    # Append image references to draft

    image_blocks = [
        f"### {name.capitalize()} View\n"

        f"![{name} projection]({url})\n"
        for name, url in views_data["views"].items()
    ]
    
    return draft_md + "\n\n## Technical Drawings\n\n" + "\n".join(image_blocks)

# Usage

final_draft = embed_cad_views(
    "/data/invention_v2.step",
    original_markdown_content
)

Source File Reference

File Responsibility Permalink
tools/shared/cad_scan.py STEP file ingestion and CADModel construction cad_scan.py
tools/shared/step_to_views.py View matrix definitions and off-screen rendering step_to_views.py
tools/shared/run_step_to_views.py CLI wrapper with JSON output serialization run_step_to_views.py
tools/shared/md_to_docx.py Document assembly consuming base-64 image payloads md_to_docx.py

Summary

  • CAD/STEP parsing begins with read-only ingestion via cad_scan.py, ensuring geometric fidelity
  • Multi-view projections are generated in step_to_views.py using predefined view matrices and off-screen OCC rendering
  • Memory-only processing through BytesIO buffers and base-64 encoding eliminates file-system side effects
  • JSON-streamed output from run_step_to_views.py enables non-blocking integration with document pipelines
  • Zero modification guarantee of source STEP files protects intellectual property documentation integrity

Frequently Asked Questions

How does the pipeline prevent accidental modification of source CAD files?

The load_step() function in cad_scan.py uses STEPControl_Reader in read-only mode and constructs an immutable CADModel dataclass. No write operations are performed on the filesystem, and the TopoDS_Shape objects held in memory are never mutated by downstream rendering code.

What distinguishes this approach from conventional CAD export workflows?

Traditional workflows require explicit file exports (PNG, SVG, PDF) to temporary directories, creating synchronization and cleanup overhead. This pipeline renders directly to memory buffers, encodes as base-64 data URLs, and streams results through stdout—eliminating temporary files entirely and reducing latency for document assembly.

Why are view projections fixed at 512×512 pixels by default?

This resolution provides sufficient detail for patent figure review while maintaining reasonable JSON payload sizes for pipeline transmission. The render_views() function accepts an optional size tuple for cases requiring higher resolution, though larger dimensions increase memory usage and serialization overhead.

Can the pipeline handle complex assemblies with multiple solid bodies?

Yes. The STEPControl_Reader.TransferRoots() call in cad_scan.py imports all root-level shapes from the STEP file. The CADModel.shape attribute holds a compound TopoDS_Shape containing all bodies, which step_to_views.py renders as a unified scene. For disaggregated views, preprocessing with OCC.Core.BRepTools shape iteration would be required upstream.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →