How the step.parts Skill Finds Off-the-Shelf Components in CAD Assemblies

The step.parts skill parses STEP file assemblies, normalizes vendor-specific part identifiers, and queries a local JSON catalogue to match geometric components with commercially available off-the-shelf parts.

The step.parts skill in the earthtojake/text-to-cad repository automates the conversion of CAD assemblies into procureable bills of materials. By analyzing geometric data and part metadata from STEP files, it bridges the gap between proprietary design formats and standardized supplier catalogues.

Parsing STEP File Assemblies to Extract Component Metadata

The skill begins by ingesting raw STEP data through the cadgen package CAD kernel. This stage focuses on decomposing the assembly hierarchy into discrete searchable entities.

Loading Geometry with cadgen.step.StepFile

In skills/step-parts/skill.py, the skill initializes the parsing pipeline by calling cadgen.step.read (exposed via cadgen.step.StepFile). This loads the entire assembly graph and prepares it for traversal.

from cadgen.step import StepFile

step = StepFile.read("models/hypercar/src/wheels.step")

The StepFile object walks the assembly hierarchy and yields individual part instances through the iter_parts() generator. Each yielded object is a cadgen.step.PartInfo instance containing the raw design data required for downstream matching.

Extracting Material and Bounding Box Attributes

For every part instance discovered, the skill extracts three critical data points:

  • Name: The vendor-specific designation embedded in the STEP entity
  • Material: The assigned material property (e.g., aluminum_6061, steel_stainless)
  • Geometric attributes: Bounding box dimensions and mass properties

These attributes are stored in the PartInfo structure and passed to the normalization layer. The bounding box data specifically enables tolerance-based matching against catalogue specifications.

Normalizing Part Identifiers for Catalogue Lookup

Raw STEP exports often contain manufacturer prefixes, revision codes, or internal project numbers that do not align with standard supplier part numbers. The skill resolves these discrepancies through a two-stage normalization process implemented in cadgen.parts.normalize_name.

Canonical Name Mapping via name_map.json

First, the skill applies a configurable translation table located at skills/step-parts/resources/name_map.json. This JSON mapping strips known vendor prefixes and converts internal codenames to canonical part identifiers recognized by the off-the-shelf catalogue.

If the extracted part name matches a key in this mapping, the skill substitutes it with the standardized equivalent before querying the database.

Fallback Fuzzy Matching by Geometry

When no explicit entry exists in the name map, the skill falls back to fuzzy matching based on geometric similarity. It compares the part’s bounding box dimensions and mass against catalogue entries to infer potential matches. This ensures that even generically named components (e.g., "bracket_left_v2") can be identified if their physical characteristics align with stocked items.

Querying the Off-the-Shelf Parts Catalogue

After normalization, the skill executes the search against the local parts database using cadgen.parts.search. The catalogue itself resides in models/parts_catalog.json and contains a comprehensive index of supplier components including part numbers, manufacturers, dimensions, pricing, and stock availability.

Filter Criteria: Part Numbers, Materials, and Tolerances

The search.catalog_lookup function applies a cascading filter to narrow results:

  1. Exact part number: Matches the normalized identifier against the catalogue SKU
  2. Material compatibility: Ensures the physical properties meet the design specification
  3. Size tolerances: Validates that bounding box dimensions fall within acceptable variance ranges
  4. Manufacturer (optional): Restricts results to preferred suppliers when specified

Ranking by Availability and Price

When multiple candidates satisfy the filter criteria, the skill ranks them according to availability (stock levels) and unit price, prioritizing components that can be procured quickly and economically. The top-ranked match is returned as the recommended off-the-shelf replacement for the original CAD part.

CLI and Python Usage Examples

The skill exposes its functionality through both a command-line interface and a Python API, as documented in skills/step-parts/references/step-parts-api.md.

Command-line usage:


# Locate off-the-shelf parts for a STEP assembly

cadgen step-parts search models/hypercar/src/wheels.step

Python usage within custom scripts:

from cadgen.step import StepFile
from cadgen.parts import search

# Load the STEP file

step = StepFile.read("models/hypercar/src/wheels.step")

# Extract parts and run the catalogue lookup

components = []
for part in step.iter_parts():
    norm_name = search.normalize_name(part.name)
    match = search.catalog_lookup(
        norm_name, 
        part.material, 
        part.bounding_box
    )
    components.append({
        "original_name": part.name,
        "catalog_match": match,
    })

print(components)

The test suite in tests/python/skills/step-parts/test_search_cli.py validates this end-to-end workflow against the bundled catalogue to ensure matching accuracy across diverse assembly types.

Summary

  • The skill parses STEP files using cadgen.step.StepFile to extract part names, materials, and bounding boxes.
  • Normalization occurs via cadgen.parts.normalize_name, utilizing skills/step-parts/resources/name_map.json for identifier translation and geometric fuzzy matching as a fallback.
  • The catalogue query against models/parts_catalog.json filters by part number, material, and dimensional tolerances, then ranks results by availability and price.
  • Users interact with the system via the cadgen step-parts search CLI command or the Python API exposed in skills/step-parts/skill.py.

Frequently Asked Questions

What input file format does the step.parts skill require?

The skill requires STEP files (ISO 10303-21) as input. The cadgen.step module handles the geometric kernel operations to parse these neutral CAD format files and extract the assembly hierarchy and part metadata.

How does the skill handle ambiguous or non-standard part names?

When a part name lacks a direct catalogue entry, the skill first consults the name_map.json configuration to translate vendor-specific codes to standard identifiers. If no mapping exists, it performs fuzzy matching based on the part’s bounding box dimensions and mass properties to infer the closest geometric equivalent.

Where is the off-the-shelf parts catalogue located?

The catalogue is stored as a JSON index at models/parts_catalog.json within the repository. This file contains structured data for each component including part numbers, manufacturers, dimensions, pricing, and current availability status.

Can I customize the search ranking criteria?

The current implementation ranks candidates primarily by availability and price as stored in the catalogue. While the step-parts-api.md documentation outlines the standard output schema, modifying the ranking weights would require adjusting the cadgen.parts.search logic in skills/step-parts/skill.py or the underlying catalogue query implementation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →