# Text-to-CAD Performance Benchmarks: A Complete Guide to the earthtojake/text-to-cad Evaluation Suite

> Explore text-to-CAD performance benchmarks using the earthtojake/text-to-cad evaluation suite. This guide details reproducible benchmarks for automated CAD generation quality and GIF visualization.

- Repository: [earthtojake/text-to-cad](https://github.com/earthtojake/text-to-cad)
- Tags: performance
- Published: 2026-08-04

---

**The earthtojake/text-to-cad repository provides ten reproducible performance benchmarks that validate CAD generation quality through automated STEP file creation and GIF visualization, with assets stored in the `benchmarks/` directory and managed via Git LFS.**

The **text-to-cad** repository by earthtojake ships a rigorous evaluation framework for assessing natural-language-to-geometry generation. These **Text-to-CAD performance benchmarks** cover diverse mechanical parts—from calibration blocks to planetary gear stages—providing visual proof of the CAD skill's capabilities. Each benchmark pairs a structured markdown description with an animated GIF output to demonstrate end-to-end geometry creation.

## Where the Benchmarks Live

### The benchmarks/ Directory and Git LFS Storage

All **Text-to-CAD performance benchmarks** reside in the top-level `benchmarks/` directory. This location stores markdown descriptions alongside animated GIF assets that visualize the generated geometry. Because the animation files are large binary assets, the repository tracks them using **Git LFS** (Large File Storage) as configured in `.gitattributes`.

To maintain a lightweight clone, `.lfsconfig` excludes these assets from default pulls. Developers can hydrate the GIFs on demand using:

```bash
git lfs pull --include="benchmarks/**"

```

### Individual Benchmark Specifications

Each benchmark follows a numbered naming convention from [`01-rectangular-calibration-block.md`](https://github.com/earthtojake/text-to-cad/blob/main/01-rectangular-calibration-block.md) through [`10-planetary-gear-stage.md`](https://github.com/earthtojake/text-to-cad/blob/main/10-planetary-gear-stage.md). These markdown files contain:

- The exact natural-language prompt sent to the CAD skill
- A description of the target geometry
- A reference to the corresponding GIF file (e.g., `benchmark_01_rectangular_calibration_block.gif`)

This **documentation-first** approach ensures reviewers can evaluate prompts and expected outputs without downloading binary assets.

## How the Benchmark Pipeline Works

The benchmark generation follows a four-stage pipeline that transforms text prompts into visual artifacts:

1. **Prompt Interpretation** – The **CAD skill** at [`skills/cad/SKILL.md`](https://github.com/earthtojake/text-to-cad/blob/main/skills/cad/SKILL.md) receives a natural-language prompt (e.g., "Create a centered 100 × 60 × 20 mm block..."). It parses the request and constructs a model using *build123d* (referenced in [`skills/cad/references/build123d-modeling.md`](https://github.com/earthtojake/text-to-cad/blob/main/skills/cad/references/build123d-modeling.md)).

2. **STEP Export** – The skill exports the generated geometry as a STEP file, the industry-standard format for CAD data exchange.

3. **Visualization** – The **CAD Viewer** ([`skills/cad-viewer/SKILL.md`](https://github.com/earthtojake/text-to-cad/blob/main/skills/cad-viewer/SKILL.md)) converts the STEP file to GLB format and renders a 360° rotation animation in a headless environment.

4. **GIF Creation** – The animation frames stitch into a GIF stored in `benchmarks/`, completing the visual documentation of the **Text-to-CAD performance benchmark**.

## Reproducing Benchmarks Locally

You can execute any benchmark locally using the command-line interface. This validates that your environment generates identical geometry to the reference implementations.

### Install the CAD Skills Library

```bash
npx skills install earthtojake/text-to-cad

```

This command installs the skill set, including the CAD skill and viewer, into your current environment.

### Run a Specific Benchmark Prompt

```bash

# Example: benchmark #3 – L-bracket

skills run cad \
  "Create an L-bracket from a base plate and rear vertical plate. \
   Add vertical base holes, horizontal back-plate holes, two triangular gussets, \
   and a filleted base/back transition."

```

This invokes the same logic found in [`skills/cad/SKILL.md`](https://github.com/earthtojake/text-to-cad/blob/main/skills/cad/SKILL.md) that generated [`benchmarks/03-l-bracket.md`](https://github.com/earthtojake/text-to-cad/blob/main/benchmarks/03-l-bracket.md), producing a STEP file and optional GLB output.

### Preview Results in the CAD Viewer

```bash
npx skills run cad-viewer --file output.step --port 4178

```

The viewer starts a local server, loads the STEP file, and generates the rotating animation. Save the output to `benchmarks/benchmark_03_l_bracket.gif` to match the repository structure.

## Testing and Validation Infrastructure

The repository ensures benchmark integrity through automated testing. The test suite at [`scripts/test/test.sh`](https://github.com/earthtojake/text-to-cad/blob/main/scripts/test/test.sh) validates that each markdown description matches its GIF output and that generated assets remain accessible.

Additionally, `viewer/src/server/vercelBlobAssetBackend.test.mjs` contains test cases verifying that requests for `benchmarks/part.step` resolve to correct URLs in the **Vercel-Blob backend**. This guarantees that benchmark files serve correctly in production deployments.

## Architectural Design Principles

### Modular Skill Isolation

Each skill lives under `skills/` and remains completely self-contained. The CAD skill does not import code from other skills; shared helpers reside in `packages/cadpy` and `packages/cadjs`. This isolation ensures **Text-to-CAD performance benchmarks** reproduce consistently across environments.

### Backend Abstraction

The viewer fetches assets through a Vercel-Blob proxy, decoupling the frontend from storage implementation. Tests in `viewer/src/server/vercelBlobAssetBackend.test.mjs` verify this abstraction layer functions correctly.

### Documentation-First Workflow

Benchmark specifications use markdown files stored in normal Git, while binary GIFs remain in LFS. This separation enables code review of prompts and descriptions without bloating repository clones.

## Summary

- The **earthtojake/text-to-cad** repository contains ten standardized benchmarks ranging from simple calibration blocks to complex planetary gear stages.
- Benchmark assets live in `benchmarks/` with markdown specifications in Git and GIF animations in Git LFS.
- The pipeline flows from [`skills/cad/SKILL.md`](https://github.com/earthtojake/text-to-cad/blob/main/skills/cad/SKILL.md) (generation) to [`skills/cad-viewer/SKILL.md`](https://github.com/earthtojake/text-to-cad/blob/main/skills/cad-viewer/SKILL.md) (visualization).
- Reproduce any benchmark locally using `skills run cad` followed by `skills run cad-viewer`.
- Automated tests in [`scripts/test/test.sh`](https://github.com/earthtojake/text-to-cad/blob/main/scripts/test/test.sh) and `vercelBlobAssetBackend.test.mjs` ensure asset integrity and backend functionality.

## Frequently Asked Questions

### How many Text-to-CAD performance benchmarks are included in the repository?

The repository includes **ten benchmarks**, numbered 01 through 10, covering geometries from rectangular calibration blocks to planetary gear stages. Each corresponds to a markdown file in `benchmarks/` (e.g., [`01-rectangular-calibration-block.md`](https://github.com/earthtojake/text-to-cad/blob/main/01-rectangular-calibration-block.md) through [`10-planetary-gear-stage.md`](https://github.com/earthtojake/text-to-cad/blob/main/10-planetary-gear-stage.md)).

### Why are the benchmark GIFs stored using Git LFS?

The animated GIF files are large binary assets that would significantly increase repository size. Git LFS filters these out of normal clones, keeping CI pipelines fast and developer clones lightweight. The configuration in `.gitattributes` tracks `benchmarks/**` files with LFS, while `.lfsconfig` prevents automatic downloads.

### How do I run a specific benchmark locally without downloading all GIFs?

Use `skills run cad` with the exact prompt text found in the benchmark's markdown file. For example, copy the prompt from [`benchmarks/03-l-bracket.md`](https://github.com/earthtojake/text-to-cad/blob/main/benchmarks/03-l-bracket.md) and pass it to the command. This generates the STEP file locally without requiring `git lfs pull`, though you will need the LFS assets if you want to compare your output GIF against the reference.

### What testing ensures the benchmarks remain accurate?

Two test suites validate benchmark integrity: [`scripts/test/test.sh`](https://github.com/earthtojake/text-to-cad/blob/main/scripts/test/test.sh) verifies that markdown descriptions match their corresponding GIF outputs, and `viewer/src/server/vercelBlobAssetBackend.test.mjs` ensures the Vercel-Blob backend correctly serves benchmark assets. These tests guarantee that the **Text-to-CAD performance benchmarks** stay synchronized with the current skill behavior.