# How to Optimize text-to-CAD for Performance: 3 Architectural Strategies

> Optimize text-to-CAD performance with three architectural strategies: content-addressed caching, deterministic builds, and a parallel daemon pool. Improve your CAD generation workflow.

- Repository: [earthtojake/text-to-cad](https://github.com/earthtojake/text-to-cad)
- Tags: performance
- Published: 2026-09-13

---

**text-to-CAD performance relies on three pillars: an immutable content-addressed cache at `~/.cache/cadgen`, deterministic builds that guarantee reproducible cache hits, and a parallel daemon pool that distributes heavy CAD kernel work across CPU cores.**

The `earthtojake/text-to-cad` repository implements a high-performance CAD generation system that transforms natural language into 3D models. To optimize text-to-CAD for performance, you must leverage its content-addressed storage layer, ensure deterministic build pipelines, and configure the parallel execution daemon for your hardware.

## The Three Performance Pillars of text-to-CAD

### Content-Addressed Object Store

The **content-addressed store** eliminates redundant computation by caching immutable objects under `~/.cache/cadgen`. When a skill calls `cadgen.step.compile`, the system first computes a content hash and checks for existing artifacts in the store.

- **Store implementation**: Defined in [`packages/cadgen/src/cadgen/store/objects.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadgen/src/cadgen/store/objects.py) and [`packages/cadgen/src/cadgen/store/index.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadgen/src/cadgen/store/index.py)
- **Atomic writes**: New artifacts are written atomically via [`packages/cadgen/src/cadgen/store/gate.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadgen/src/cadgen/store/gate.py) to prevent corruption during parallel access
- **Cache hits**: Identical inputs reuse existing geometry instantly, skipping kernel recomputation entirely

### Deterministic Build Pipeline

**Deterministic builds** ensure that identical source code always produces byte-identical outputs, which is essential for reliable cache hits. This architecture is implemented in [`packages/cadgen/src/cadgen/_internal/generation.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadgen/src/cadgen/_internal/generation.py) and [`authoring.py`](https://github.com/earthtojake/text-to-cad/blob/main/authoring.py).

- Decorator arguments affect only side-car metadata, never geometry itself
- Design law 16 dictates that metadata changes should not invalidate geometry cache entries
- Guarantees that the same source produces the same bytes across different runs and environments

### Parallel Daemon Pool

The **daemon pool** handles CPU-intensive operations like STEP compilation and tessellation across multiple workers. Located in `packages/cadgen/src/cadgen/daemon/` (including [`executors.py`](https://github.com/earthtojake/text-to-cad/blob/main/executors.py), broker, and pool modules), this background system hides latency by distributing work.

- Workers, spares, and a job ledger manage heavy CAD kernel operations
- Parallel execution utilizes all available CPU cores for geometry generation
- The pool automatically spins up jobs on cache misses, writing results to the store upon completion

## How the Cache Pipeline Works

When you trigger a build, the system follows this execution path:

1. **Hash calculation**: The input source is hashed to generate a content address
2. **Cache lookup**: The system queries the store for an existing artifact matching that hash
3. **Fast path**: On cache hit, the stored STEP/STL file and its side-car metadata return instantly without kernel execution
4. **Compute path**: On cache miss, the daemon pool assigns the job to an available worker
5. **Storage**: The kernel writes the result atomically to the store via [`gate.py`](https://github.com/earthtojake/text-to-cad/blob/main/gate.py), ensuring subsequent calls hit the cache

## Environment Variables for Performance Tuning

Control the optimization behavior through these environment variables:

- **`CADGEN_CACHE_DIR`**: Overrides the default `~/.cache/cadgen` location. Point this to a fast local SSD for significant I/O gains. Example: `export CADGEN_CACHE_DIR=/mnt/ssd/cadgen-cache` as demonstrated in [`tests/python/support/cad_test_roots.py`](https://github.com/earthtojake/text-to-cad/blob/main/tests/python/support/cad_test_roots.py).

- **`CADGEN_DAEMON_WORKERS`**: Sets the number of parallel workers the daemon spawns. Match this to your physical CPU core count for optimal throughput (e.g., `export CADGEN_DAEMON_WORKERS=8` for an 8-core machine).

- **`CADGEN_DAEMON_TIMEOUT`**: Maximum seconds a worker may block before being culled. Increase this for complex geometries: `export CADGEN_DAEMON_TIMEOUT=600`.

- **`CADGEN_DISABLE_CACHE`**: Set to `1` to completely disable the store. Useful for debugging geometry generation issues, but eliminates all caching benefits and dramatically slows builds.

## Optimization Best Practices

- **Persist the cache on fast storage**: Keep `CADGEN_CACHE_DIR` on a local SSD and version-control its location in a `.env` file to ensure CI pipelines and repeated local runs reuse artifacts.

- **Minimize side-car churn**: Keep decorator arguments minimal and stable. Unnecessary metadata changes break cache hits even when geometry remains unchanged.

- **Right-size parallelism**: Set `CADGEN_DAEMON_WORKERS` to match physical CPU cores; monitor actual utilization via `cadgen daemon status` to verify the pool saturates available compute without oversubscribing.

- **Maintain determinism**: Avoid non-deterministic randomness in model code. Use fixed random seeds or pure functions to ensure the deterministic pipeline in [`_internal/generation.py`](https://github.com/earthtojake/text-to-cad/blob/main/_internal/generation.py) produces consistent hashes.

## Implementation Examples

Configure a high-speed SSD cache:

```python
import os

os.environ["CADGEN_CACHE_DIR"] = "/mnt/ssd/cadgen-cache"

# Subsequent cadgen calls will reuse cached artifacts from SSD

```

Run heavy STEP compilation with automatic parallelization:

```python
import cadgen.step as step

# Assuming my_model.py defines a @step model

step.compile("my_model.py", out="part.step")  # Fast on cache hit

# The daemon automatically spawns workers if a cache miss occurs

```

Tune the daemon for an 8-core workstation:

```bash
export CADGEN_DAEMON_WORKERS=8
export CADGEN_DAEMON_TIMEOUT=600   # 10 minutes max per job

cadgen viewer start  # Starts the viewer and tuned daemon

```

Explicitly bypass the cache for debugging:

```python
import os

os.environ["CADGEN_DISABLE_CACHE"] = "1"
import cadgen.step as step
step.compile("my_model.py", out="fresh.step")

```

## Summary

- The **content-addressed store** in `packages/cadgen/src/cadgen/store/` eliminates redundant computation by reusing immutable cached objects stored under `~/.cache/cadgen`.
- **Deterministic builds** in [`_internal/generation.py`](https://github.com/earthtojake/text-to-cad/blob/main/_internal/generation.py) ensure identical inputs produce byte-identical outputs, maximizing cache hit rates across runs.
- The **daemon pool** in [`cadgen/daemon/executors.py`](https://github.com/earthtojake/text-to-cad/blob/main/cadgen/daemon/executors.py) parallelizes heavy kernel operations like STEP compilation across configurable workers.
- Environment variables `CADGEN_CACHE_DIR` and `CADGEN_DAEMON_WORKERS` provide direct control over storage location and parallelism levels.
- Avoid `CADGEN_DISABLE_CACHE` in production environments, as it forces full recomputation of all geometry and eliminates performance gains.

## Frequently Asked Questions

### Where does text-to-CAD store its cache by default?

The system stores immutable content-addressed objects in `~/.cache/cadgen` by default. You can override this location by setting the `CADGEN_CACHE_DIR` environment variable to any fast storage path, as demonstrated in [`tests/python/support/cad_test_roots.py`](https://github.com/earthtojake/text-to-cad/blob/main/tests/python/support/cad_test_roots.py).

### How many daemon workers should I configure for optimal performance?

Set `CADGEN_DAEMON_WORKERS` to match your machine's physical CPU core count. For an 8-core machine, use `export CADGEN_DAEMON_WORKERS=8`. Monitor actual CPU utilization via `cadgen daemon status` to verify the pool is saturating available cores without causing system oversubscription.

### Why are my builds not hitting the cache despite identical code?

Cache misses typically indicate non-determinism in your build pipeline or excessive churn in decorator side-car metadata. Ensure your model code uses fixed random seeds and that decorator arguments remain stable between runs, as the deterministic pipeline requires byte-identical inputs to generate matching content hashes.

### Can I disable caching for debugging purposes?

Yes, set `export CADGEN_DISABLE_CACHE=1` to bypass the content-addressed store. This forces the daemon to execute every kernel operation fresh, which is useful for debugging geometry generation issues but dramatically reduces performance. Remove this variable for normal operations to restore caching benefits.