CadPy Catalog System for Indexing CAD Sources: A Complete Technical Guide

The CadPy catalog system is a pure-Python registry that automatically discovers, validates, and indexes Python generator scripts and STEP files to create a unified CAD artifact index with O(1) lookup performance.

The CadPy catalog system for indexing CAD sources serves as the central nervous system of the earthtojake/text-to-cad repository. This lightweight framework eliminates manual bookkeeping by scanning the filesystem for manufacturable components and enforcing strict uniqueness constraints. Because the catalog is implemented entirely in Python without external database dependencies, it can be recomputed on-the-fly during CI pipelines or local development workflows.

Core Architecture and Repository Roots

In packages/cadpy/src/cadpy/catalog.py, the catalog initializes by establishing filesystem anchors that define the search boundaries for the discovery engine.

Repository Root Definitions

The module resolves REPO_ROOT via Path.cwd().resolve() and sets CAD_ROOT to match this location. These constants determine where the catalog begins its recursive scan for CAD artifacts.

The Immutable CadSource Dataclass

Every indexed entry becomes a CadSource instance—a frozen @dataclass defined in catalog.py (lines 68-88) that captures metadata including file paths, generator kind, tolerance settings, and reference identifiers. Because the class is frozen, records are hashable and safe for use as dictionary keys, ensuring integrity during downstream processing.

Automatic Discovery and Validation

The iter_cad_sources() function serves as the main entry point, orchestrating two distinct discovery strategies to build the complete index.

Python Generator Script Discovery

When _iter_python_sources() scans the repository, it identifies *.py files containing a gen_step function. The helper _read_python_source() then parses the script to extract generator metadata via parse_generator_metadata. For scripts that generate DXF without STEP geometry, the system supports an allow_dxf_only=True flag to include these sources in the catalog.

Standalone STEP File Indexing

For legacy or external geometry, _iter_step_sources() locates *.step and *.stp files. The _read_step_source() function constructs a CadSource record for these raw files, treating them as self-contained artifacts without generating scripts.

Duplicate Detection and Conflict Resolution

During indexing, the system enforces uniqueness across four critical dimensions: cad_ref, source_ref, STEP file paths, and generated artifact paths. If any collision occurs, the catalog raises a CadSourceError immediately, preventing ambiguous references from entering the index and ensuring a single source of truth.

Query Interface and Reference Normalization

Once indexed, the catalog provides fast lookup utilities that normalize references before querying.

Bulk Lookup Dictionaries

The source_by_cad_ref() function returns a complete dictionary mapping normalized cad_ref strings to their CadSource objects. This enables constant-time access for bulk operations across the entire repository.

Individual Source Resolution

For targeted queries, find_source_by_cad_ref() and find_source_by_source_ref() provide convenient lookups. These functions automatically normalize input strings using normalize_cad_ref and normalize_source_ref imported from packages/cadpy/src/cadpy/metadata.py, ensuring consistent key matching regardless of formatting variations.

Generated Artifact Path Resolution

The catalog computes derived file locations for downstream consumers. The explorer_artifact_path_for_step_path() helper calculates hidden and legacy paths for GLB meshes, topology JSON files, and other explorer artifacts based on the source STEP file location.

Practical Usage Examples

These snippets demonstrate the typical workflow: discover → query → derive artifact locations.


# List every CAD source discovered in the repository

from cadpy.catalog import iter_cad_sources

for src in iter_cad_sources():
    print(f"{src.cad_ref=:<12} {src.kind=:<8} {src.source_path}")

# Find a specific source by its human-readable CAD reference

from cadpy.catalog import find_source_by_cad_ref

cad_ref = "my_robot/base_plate"
src = find_source_by_cad_ref(cad_ref)
if src:
    print(f"Found STEP at {src.step_path}")
else:
    print("No such CAD reference")

# Resolve the explorer GLB artifact for a generated STEP file

from pathlib import Path
from cadpy.catalog import explorer_artifact_path_for_step_path

step_path = Path("models/assembly/arm.step")
glb_path = explorer_artifact_path_for_step_path(step_path, ".glb")
print(f"GLB artifact will be stored at {glb_path}")

Summary

  • The CadPy catalog in packages/cadpy/src/cadpy/catalog.py provides automatic discovery of Python generators and STEP files.
  • Uniqueness enforcement prevents duplicate cad_ref, source_ref, or path entries via CadSourceError.
  • Immutable records use the frozen CadSource dataclass to ensure hashable, consistent metadata.
  • O(1) lookups are available through source_by_cad_ref(), find_source_by_cad_ref(), and find_source_by_source_ref().
  • Reference normalization via metadata.py ensures stable keys regardless of input formatting.
  • Artifact resolution helpers calculate derived paths for GLB and JSON files without external database queries.

Frequently Asked Questions

What file types does the CadPy catalog system index?

The catalog indexes two primary source types: Python generator scripts (*.py) that declare a gen_step function, and standalone STEP files (*.step or *.stp). Python sources may optionally be DXF-only when the allow_dxf_only flag is enabled during parsing.

How does the catalog prevent duplicate CAD references?

During the indexing process in iter_cad_sources(), the system checks for conflicts across cad_ref, source_ref, STEP paths, and generated artifact paths. If any duplicate is detected, the catalog raises a CadSourceError immediately, ensuring every identifier in the repository remains unique.

Where does the catalog store its index data?

The catalog does not use an external database or persistent storage. Instead, it recomputes the index on-the-fly by scanning the filesystem defined by REPO_ROOT and CAD_ROOT. This design allows the index to remain fresh in CI pipelines and eliminates synchronization issues between disk and database states.

How are generated artifact paths like GLB files determined?

The catalog provides helper functions such as explorer_artifact_path_for_step_path() that compute derived file locations based on the source STEP file path. These functions follow convention-based rules to determine where GLB meshes, topology JSON files, and other explorer artifacts should be stored or retrieved.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →