Text-to-CAD Performance Benchmarks: A Complete Guide to the earthtojake/text-to-cad Evaluation Suite

The earthtojake/text-to-cad repository provides ten reproducible performance benchmarks that validate CAD generation quality through automated STEP file creation and GIF visualization, with assets stored in the benchmarks/ directory and managed via Git LFS.

The text-to-cad repository by earthtojake ships a rigorous evaluation framework for assessing natural-language-to-geometry generation. These Text-to-CAD performance benchmarks cover diverse mechanical parts—from calibration blocks to planetary gear stages—providing visual proof of the CAD skill's capabilities. Each benchmark pairs a structured markdown description with an animated GIF output to demonstrate end-to-end geometry creation.

Where the Benchmarks Live

The benchmarks/ Directory and Git LFS Storage

All Text-to-CAD performance benchmarks reside in the top-level benchmarks/ directory. This location stores markdown descriptions alongside animated GIF assets that visualize the generated geometry. Because the animation files are large binary assets, the repository tracks them using Git LFS (Large File Storage) as configured in .gitattributes.

To maintain a lightweight clone, .lfsconfig excludes these assets from default pulls. Developers can hydrate the GIFs on demand using:

git lfs pull --include="benchmarks/**"

Individual Benchmark Specifications

Each benchmark follows a numbered naming convention from 01-rectangular-calibration-block.md through 10-planetary-gear-stage.md. These markdown files contain:

  • The exact natural-language prompt sent to the CAD skill
  • A description of the target geometry
  • A reference to the corresponding GIF file (e.g., benchmark_01_rectangular_calibration_block.gif)

This documentation-first approach ensures reviewers can evaluate prompts and expected outputs without downloading binary assets.

How the Benchmark Pipeline Works

The benchmark generation follows a four-stage pipeline that transforms text prompts into visual artifacts:

  1. Prompt Interpretation – The CAD skill at skills/cad/SKILL.md receives a natural-language prompt (e.g., "Create a centered 100 × 60 × 20 mm block..."). It parses the request and constructs a model using build123d (referenced in skills/cad/references/build123d-modeling.md).

  2. STEP Export – The skill exports the generated geometry as a STEP file, the industry-standard format for CAD data exchange.

  3. Visualization – The CAD Viewer (skills/cad-viewer/SKILL.md) converts the STEP file to GLB format and renders a 360° rotation animation in a headless environment.

  4. GIF Creation – The animation frames stitch into a GIF stored in benchmarks/, completing the visual documentation of the Text-to-CAD performance benchmark.

Reproducing Benchmarks Locally

You can execute any benchmark locally using the command-line interface. This validates that your environment generates identical geometry to the reference implementations.

Install the CAD Skills Library

npx skills install earthtojake/text-to-cad

This command installs the skill set, including the CAD skill and viewer, into your current environment.

Run a Specific Benchmark Prompt


# Example: benchmark #3 – L-bracket

skills run cad \
  "Create an L-bracket from a base plate and rear vertical plate. \
   Add vertical base holes, horizontal back-plate holes, two triangular gussets, \
   and a filleted base/back transition."

This invokes the same logic found in skills/cad/SKILL.md that generated benchmarks/03-l-bracket.md, producing a STEP file and optional GLB output.

Preview Results in the CAD Viewer

npx skills run cad-viewer --file output.step --port 4178

The viewer starts a local server, loads the STEP file, and generates the rotating animation. Save the output to benchmarks/benchmark_03_l_bracket.gif to match the repository structure.

Testing and Validation Infrastructure

The repository ensures benchmark integrity through automated testing. The test suite at scripts/test/test.sh validates that each markdown description matches its GIF output and that generated assets remain accessible.

Additionally, viewer/src/server/vercelBlobAssetBackend.test.mjs contains test cases verifying that requests for benchmarks/part.step resolve to correct URLs in the Vercel-Blob backend. This guarantees that benchmark files serve correctly in production deployments.

Architectural Design Principles

Modular Skill Isolation

Each skill lives under skills/ and remains completely self-contained. The CAD skill does not import code from other skills; shared helpers reside in packages/cadpy and packages/cadjs. This isolation ensures Text-to-CAD performance benchmarks reproduce consistently across environments.

Backend Abstraction

The viewer fetches assets through a Vercel-Blob proxy, decoupling the frontend from storage implementation. Tests in viewer/src/server/vercelBlobAssetBackend.test.mjs verify this abstraction layer functions correctly.

Documentation-First Workflow

Benchmark specifications use markdown files stored in normal Git, while binary GIFs remain in LFS. This separation enables code review of prompts and descriptions without bloating repository clones.

Summary

  • The earthtojake/text-to-cad repository contains ten standardized benchmarks ranging from simple calibration blocks to complex planetary gear stages.
  • Benchmark assets live in benchmarks/ with markdown specifications in Git and GIF animations in Git LFS.
  • The pipeline flows from skills/cad/SKILL.md (generation) to skills/cad-viewer/SKILL.md (visualization).
  • Reproduce any benchmark locally using skills run cad followed by skills run cad-viewer.
  • Automated tests in scripts/test/test.sh and vercelBlobAssetBackend.test.mjs ensure asset integrity and backend functionality.

Frequently Asked Questions

How many Text-to-CAD performance benchmarks are included in the repository?

The repository includes ten benchmarks, numbered 01 through 10, covering geometries from rectangular calibration blocks to planetary gear stages. Each corresponds to a markdown file in benchmarks/ (e.g., 01-rectangular-calibration-block.md through 10-planetary-gear-stage.md).

Why are the benchmark GIFs stored using Git LFS?

The animated GIF files are large binary assets that would significantly increase repository size. Git LFS filters these out of normal clones, keeping CI pipelines fast and developer clones lightweight. The configuration in .gitattributes tracks benchmarks/** files with LFS, while .lfsconfig prevents automatic downloads.

How do I run a specific benchmark locally without downloading all GIFs?

Use skills run cad with the exact prompt text found in the benchmark's markdown file. For example, copy the prompt from benchmarks/03-l-bracket.md and pass it to the command. This generates the STEP file locally without requiring git lfs pull, though you will need the LFS assets if you want to compare your output GIF against the reference.

What testing ensures the benchmarks remain accurate?

Two test suites validate benchmark integrity: scripts/test/test.sh verifies that markdown descriptions match their corresponding GIF outputs, and viewer/src/server/vercelBlobAssetBackend.test.mjs ensures the Vercel-Blob backend correctly serves benchmark assets. These tests guarantee that the Text-to-CAD performance benchmarks stay synchronized with the current skill behavior.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →