# How the step.parts Skill Finds Off-the-Shelf Components in CAD Assemblies

> Discover how the step.parts skill efficiently finds off-the-shelf components in CAD assemblies by parsing STEP files and querying a JSON catalog. Simplify your design process.

- Repository: [earthtojake/text-to-cad](https://github.com/earthtojake/text-to-cad)
- Tags: how-to-guide
- Published: 2026-09-11

---

**The step.parts skill parses STEP file assemblies, normalizes vendor-specific part identifiers, and queries a local JSON catalogue to match geometric components with commercially available off-the-shelf parts.**

The `step.parts` skill in the `earthtojake/text-to-cad` repository automates the conversion of CAD assemblies into procureable bills of materials. By analyzing geometric data and part metadata from STEP files, it bridges the gap between proprietary design formats and standardized supplier catalogues.

## Parsing STEP File Assemblies to Extract Component Metadata

The skill begins by ingesting raw STEP data through the `cadgen` package CAD kernel. This stage focuses on decomposing the assembly hierarchy into discrete searchable entities.

### Loading Geometry with cadgen.step.StepFile

In [`skills/step-parts/skill.py`](https://github.com/earthtojake/text-to-cad/blob/main/skills/step-parts/skill.py), the skill initializes the parsing pipeline by calling `cadgen.step.read` (exposed via `cadgen.step.StepFile`). This loads the entire assembly graph and prepares it for traversal.

```python
from cadgen.step import StepFile

step = StepFile.read("models/hypercar/src/wheels.step")

```

The `StepFile` object walks the assembly hierarchy and yields individual part instances through the `iter_parts()` generator. Each yielded object is a `cadgen.step.PartInfo` instance containing the raw design data required for downstream matching.

### Extracting Material and Bounding Box Attributes

For every part instance discovered, the skill extracts three critical data points:

- **Name**: The vendor-specific designation embedded in the STEP entity
- **Material**: The assigned material property (e.g., *aluminum_6061*, *steel_stainless*)
- **Geometric attributes**: Bounding box dimensions and mass properties

These attributes are stored in the `PartInfo` structure and passed to the normalization layer. The bounding box data specifically enables tolerance-based matching against catalogue specifications.

## Normalizing Part Identifiers for Catalogue Lookup

Raw STEP exports often contain manufacturer prefixes, revision codes, or internal project numbers that do not align with standard supplier part numbers. The skill resolves these discrepancies through a two-stage normalization process implemented in `cadgen.parts.normalize_name`.

### Canonical Name Mapping via name_map.json

First, the skill applies a configurable translation table located at [`skills/step-parts/resources/name_map.json`](https://github.com/earthtojake/text-to-cad/blob/main/skills/step-parts/resources/name_map.json). This JSON mapping strips known vendor prefixes and converts internal codenames to canonical part identifiers recognized by the off-the-shelf catalogue.

If the extracted part name matches a key in this mapping, the skill substitutes it with the standardized equivalent before querying the database.

### Fallback Fuzzy Matching by Geometry

When no explicit entry exists in the name map, the skill falls back to fuzzy matching based on geometric similarity. It compares the part’s bounding box dimensions and mass against catalogue entries to infer potential matches. This ensures that even generically named components (e.g., *"bracket_left_v2"*) can be identified if their physical characteristics align with stocked items.

## Querying the Off-the-Shelf Parts Catalogue

After normalization, the skill executes the search against the local parts database using `cadgen.parts.search`. The catalogue itself resides in [`models/parts_catalog.json`](https://github.com/earthtojake/text-to-cad/blob/main/models/parts_catalog.json) and contains a comprehensive index of supplier components including part numbers, manufacturers, dimensions, pricing, and stock availability.

### Filter Criteria: Part Numbers, Materials, and Tolerances

The `search.catalog_lookup` function applies a cascading filter to narrow results:

1. **Exact part number**: Matches the normalized identifier against the catalogue SKU
2. **Material compatibility**: Ensures the physical properties meet the design specification
3. **Size tolerances**: Validates that bounding box dimensions fall within acceptable variance ranges
4. **Manufacturer** (optional): Restricts results to preferred suppliers when specified

### Ranking by Availability and Price

When multiple candidates satisfy the filter criteria, the skill ranks them according to **availability** (stock levels) and **unit price**, prioritizing components that can be procured quickly and economically. The top-ranked match is returned as the recommended off-the-shelf replacement for the original CAD part.

## CLI and Python Usage Examples

The skill exposes its functionality through both a command-line interface and a Python API, as documented in [`skills/step-parts/references/step-parts-api.md`](https://github.com/earthtojake/text-to-cad/blob/main/skills/step-parts/references/step-parts-api.md).

**Command-line usage:**

```bash

# Locate off-the-shelf parts for a STEP assembly

cadgen step-parts search models/hypercar/src/wheels.step

```

**Python usage within custom scripts:**

```python
from cadgen.step import StepFile
from cadgen.parts import search

# Load the STEP file

step = StepFile.read("models/hypercar/src/wheels.step")

# Extract parts and run the catalogue lookup

components = []
for part in step.iter_parts():
    norm_name = search.normalize_name(part.name)
    match = search.catalog_lookup(
        norm_name, 
        part.material, 
        part.bounding_box
    )
    components.append({
        "original_name": part.name,
        "catalog_match": match,
    })

print(components)

```

The test suite in [`tests/python/skills/step-parts/test_search_cli.py`](https://github.com/earthtojake/text-to-cad/blob/main/tests/python/skills/step-parts/test_search_cli.py) validates this end-to-end workflow against the bundled catalogue to ensure matching accuracy across diverse assembly types.

## Summary

- The skill parses STEP files using `cadgen.step.StepFile` to extract part names, materials, and bounding boxes.
- Normalization occurs via `cadgen.parts.normalize_name`, utilizing [`skills/step-parts/resources/name_map.json`](https://github.com/earthtojake/text-to-cad/blob/main/skills/step-parts/resources/name_map.json) for identifier translation and geometric fuzzy matching as a fallback.
- The catalogue query against [`models/parts_catalog.json`](https://github.com/earthtojake/text-to-cad/blob/main/models/parts_catalog.json) filters by part number, material, and dimensional tolerances, then ranks results by availability and price.
- Users interact with the system via the `cadgen step-parts search` CLI command or the Python API exposed in [`skills/step-parts/skill.py`](https://github.com/earthtojake/text-to-cad/blob/main/skills/step-parts/skill.py).

## Frequently Asked Questions

### What input file format does the step.parts skill require?

The skill requires **STEP files** (ISO 10303-21) as input. The `cadgen.step` module handles the geometric kernel operations to parse these neutral CAD format files and extract the assembly hierarchy and part metadata.

### How does the skill handle ambiguous or non-standard part names?

When a part name lacks a direct catalogue entry, the skill first consults the [`name_map.json`](https://github.com/earthtojake/text-to-cad/blob/main/name_map.json) configuration to translate vendor-specific codes to standard identifiers. If no mapping exists, it performs fuzzy matching based on the part’s bounding box dimensions and mass properties to infer the closest geometric equivalent.

### Where is the off-the-shelf parts catalogue located?

The catalogue is stored as a JSON index at [`models/parts_catalog.json`](https://github.com/earthtojake/text-to-cad/blob/main/models/parts_catalog.json) within the repository. This file contains structured data for each component including part numbers, manufacturers, dimensions, pricing, and current availability status.

### Can I customize the search ranking criteria?

The current implementation ranks candidates primarily by **availability** and **price** as stored in the catalogue. While the [`step-parts-api.md`](https://github.com/earthtojake/text-to-cad/blob/main/step-parts-api.md) documentation outlines the standard output schema, modifying the ranking weights would require adjusting the `cadgen.parts.search` logic in [`skills/step-parts/skill.py`](https://github.com/earthtojake/text-to-cad/blob/main/skills/step-parts/skill.py) or the underlying catalogue query implementation.