# CadPy Catalog System for Indexing CAD Sources: A Complete Technical Guide

> Explore the CadPy catalog system, a Python registry that indexes CAD generator scripts and STEP files for fast O(1) lookups. Discover a unified CAD artifact index.

- Repository: [earthtojake/text-to-cad](https://github.com/earthtojake/text-to-cad)
- Tags: how-to-guide
- Published: 2026-08-01

---

**The CadPy catalog system is a pure-Python registry that automatically discovers, validates, and indexes Python generator scripts and STEP files to create a unified CAD artifact index with O(1) lookup performance.**

The CadPy catalog system for indexing CAD sources serves as the central nervous system of the `earthtojake/text-to-cad` repository. This lightweight framework eliminates manual bookkeeping by scanning the filesystem for manufacturable components and enforcing strict uniqueness constraints. Because the catalog is implemented entirely in Python without external database dependencies, it can be recomputed on-the-fly during CI pipelines or local development workflows.

## Core Architecture and Repository Roots

In [`packages/cadpy/src/cadpy/catalog.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadpy/src/cadpy/catalog.py), the catalog initializes by establishing filesystem anchors that define the search boundaries for the discovery engine.

### Repository Root Definitions

The module resolves `REPO_ROOT` via `Path.cwd().resolve()` and sets `CAD_ROOT` to match this location. These constants determine where the catalog begins its recursive scan for CAD artifacts.

### The Immutable CadSource Dataclass

Every indexed entry becomes a `CadSource` instance—a frozen `@dataclass` defined in [`catalog.py`](https://github.com/earthtojake/text-to-cad/blob/main/catalog.py) (lines 68-88) that captures metadata including file paths, generator kind, tolerance settings, and reference identifiers. Because the class is frozen, records are hashable and safe for use as dictionary keys, ensuring integrity during downstream processing.

## Automatic Discovery and Validation

The `iter_cad_sources()` function serves as the main entry point, orchestrating two distinct discovery strategies to build the complete index.

### Python Generator Script Discovery

When `_iter_python_sources()` scans the repository, it identifies `*.py` files containing a `gen_step` function. The helper `_read_python_source()` then parses the script to extract generator metadata via `parse_generator_metadata`. For scripts that generate DXF without STEP geometry, the system supports an `allow_dxf_only=True` flag to include these sources in the catalog.

### Standalone STEP File Indexing

For legacy or external geometry, `_iter_step_sources()` locates `*.step` and `*.stp` files. The `_read_step_source()` function constructs a `CadSource` record for these raw files, treating them as self-contained artifacts without generating scripts.

### Duplicate Detection and Conflict Resolution

During indexing, the system enforces uniqueness across four critical dimensions: `cad_ref`, `source_ref`, STEP file paths, and generated artifact paths. If any collision occurs, the catalog raises a `CadSourceError` immediately, preventing ambiguous references from entering the index and ensuring a single source of truth.

## Query Interface and Reference Normalization

Once indexed, the catalog provides fast lookup utilities that normalize references before querying.

### Bulk Lookup Dictionaries

The `source_by_cad_ref()` function returns a complete dictionary mapping normalized `cad_ref` strings to their `CadSource` objects. This enables constant-time access for bulk operations across the entire repository.

### Individual Source Resolution

For targeted queries, `find_source_by_cad_ref()` and `find_source_by_source_ref()` provide convenient lookups. These functions automatically normalize input strings using `normalize_cad_ref` and `normalize_source_ref` imported from [`packages/cadpy/src/cadpy/metadata.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadpy/src/cadpy/metadata.py), ensuring consistent key matching regardless of formatting variations.

## Generated Artifact Path Resolution

The catalog computes derived file locations for downstream consumers. The `explorer_artifact_path_for_step_path()` helper calculates hidden and legacy paths for GLB meshes, topology JSON files, and other explorer artifacts based on the source STEP file location.

## Practical Usage Examples

These snippets demonstrate the typical workflow: **discover → query → derive artifact locations**.

```python

# List every CAD source discovered in the repository

from cadpy.catalog import iter_cad_sources

for src in iter_cad_sources():
    print(f"{src.cad_ref=:<12} {src.kind=:<8} {src.source_path}")

```

```python

# Find a specific source by its human-readable CAD reference

from cadpy.catalog import find_source_by_cad_ref

cad_ref = "my_robot/base_plate"
src = find_source_by_cad_ref(cad_ref)
if src:
    print(f"Found STEP at {src.step_path}")
else:
    print("No such CAD reference")

```

```python

# Resolve the explorer GLB artifact for a generated STEP file

from pathlib import Path
from cadpy.catalog import explorer_artifact_path_for_step_path

step_path = Path("models/assembly/arm.step")
glb_path = explorer_artifact_path_for_step_path(step_path, ".glb")
print(f"GLB artifact will be stored at {glb_path}")

```

## Summary

- The **CadPy catalog** in [`packages/cadpy/src/cadpy/catalog.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadpy/src/cadpy/catalog.py) provides automatic discovery of Python generators and STEP files.
- **Uniqueness enforcement** prevents duplicate `cad_ref`, `source_ref`, or path entries via `CadSourceError`.
- **Immutable records** use the frozen `CadSource` dataclass to ensure hashable, consistent metadata.
- **O(1) lookups** are available through `source_by_cad_ref()`, `find_source_by_cad_ref()`, and `find_source_by_source_ref()`.
- **Reference normalization** via [`metadata.py`](https://github.com/earthtojake/text-to-cad/blob/main/metadata.py) ensures stable keys regardless of input formatting.
- **Artifact resolution** helpers calculate derived paths for GLB and JSON files without external database queries.

## Frequently Asked Questions

### What file types does the CadPy catalog system index?

The catalog indexes two primary source types: Python generator scripts (`*.py`) that declare a `gen_step` function, and standalone STEP files (`*.step` or `*.stp`). Python sources may optionally be DXF-only when the `allow_dxf_only` flag is enabled during parsing.

### How does the catalog prevent duplicate CAD references?

During the indexing process in `iter_cad_sources()`, the system checks for conflicts across `cad_ref`, `source_ref`, STEP paths, and generated artifact paths. If any duplicate is detected, the catalog raises a `CadSourceError` immediately, ensuring every identifier in the repository remains unique.

### Where does the catalog store its index data?

The catalog does not use an external database or persistent storage. Instead, it recomputes the index on-the-fly by scanning the filesystem defined by `REPO_ROOT` and `CAD_ROOT`. This design allows the index to remain fresh in CI pipelines and eliminates synchronization issues between disk and database states.

### How are generated artifact paths like GLB files determined?

The catalog provides helper functions such as `explorer_artifact_path_for_step_path()` that compute derived file locations based on the source STEP file path. These functions follow convention-based rules to determine where GLB meshes, topology JSON files, and other explorer artifacts should be stored or retrieved.