How to Optimize text-to-CAD for Performance: 3 Architectural Strategies
text-to-CAD performance relies on three pillars: an immutable content-addressed cache at ~/.cache/cadgen, deterministic builds that guarantee reproducible cache hits, and a parallel daemon pool that distributes heavy CAD kernel work across CPU cores.
The earthtojake/text-to-cad repository implements a high-performance CAD generation system that transforms natural language into 3D models. To optimize text-to-CAD for performance, you must leverage its content-addressed storage layer, ensure deterministic build pipelines, and configure the parallel execution daemon for your hardware.
The Three Performance Pillars of text-to-CAD
Content-Addressed Object Store
The content-addressed store eliminates redundant computation by caching immutable objects under ~/.cache/cadgen. When a skill calls cadgen.step.compile, the system first computes a content hash and checks for existing artifacts in the store.
- Store implementation: Defined in
packages/cadgen/src/cadgen/store/objects.pyandpackages/cadgen/src/cadgen/store/index.py - Atomic writes: New artifacts are written atomically via
packages/cadgen/src/cadgen/store/gate.pyto prevent corruption during parallel access - Cache hits: Identical inputs reuse existing geometry instantly, skipping kernel recomputation entirely
Deterministic Build Pipeline
Deterministic builds ensure that identical source code always produces byte-identical outputs, which is essential for reliable cache hits. This architecture is implemented in packages/cadgen/src/cadgen/_internal/generation.py and authoring.py.
- Decorator arguments affect only side-car metadata, never geometry itself
- Design law 16 dictates that metadata changes should not invalidate geometry cache entries
- Guarantees that the same source produces the same bytes across different runs and environments
Parallel Daemon Pool
The daemon pool handles CPU-intensive operations like STEP compilation and tessellation across multiple workers. Located in packages/cadgen/src/cadgen/daemon/ (including executors.py, broker, and pool modules), this background system hides latency by distributing work.
- Workers, spares, and a job ledger manage heavy CAD kernel operations
- Parallel execution utilizes all available CPU cores for geometry generation
- The pool automatically spins up jobs on cache misses, writing results to the store upon completion
How the Cache Pipeline Works
When you trigger a build, the system follows this execution path:
- Hash calculation: The input source is hashed to generate a content address
- Cache lookup: The system queries the store for an existing artifact matching that hash
- Fast path: On cache hit, the stored STEP/STL file and its side-car metadata return instantly without kernel execution
- Compute path: On cache miss, the daemon pool assigns the job to an available worker
- Storage: The kernel writes the result atomically to the store via
gate.py, ensuring subsequent calls hit the cache
Environment Variables for Performance Tuning
Control the optimization behavior through these environment variables:
-
CADGEN_CACHE_DIR: Overrides the default~/.cache/cadgenlocation. Point this to a fast local SSD for significant I/O gains. Example:export CADGEN_CACHE_DIR=/mnt/ssd/cadgen-cacheas demonstrated intests/python/support/cad_test_roots.py. -
CADGEN_DAEMON_WORKERS: Sets the number of parallel workers the daemon spawns. Match this to your physical CPU core count for optimal throughput (e.g.,export CADGEN_DAEMON_WORKERS=8for an 8-core machine). -
CADGEN_DAEMON_TIMEOUT: Maximum seconds a worker may block before being culled. Increase this for complex geometries:export CADGEN_DAEMON_TIMEOUT=600. -
CADGEN_DISABLE_CACHE: Set to1to completely disable the store. Useful for debugging geometry generation issues, but eliminates all caching benefits and dramatically slows builds.
Optimization Best Practices
-
Persist the cache on fast storage: Keep
CADGEN_CACHE_DIRon a local SSD and version-control its location in a.envfile to ensure CI pipelines and repeated local runs reuse artifacts. -
Minimize side-car churn: Keep decorator arguments minimal and stable. Unnecessary metadata changes break cache hits even when geometry remains unchanged.
-
Right-size parallelism: Set
CADGEN_DAEMON_WORKERSto match physical CPU cores; monitor actual utilization viacadgen daemon statusto verify the pool saturates available compute without oversubscribing. -
Maintain determinism: Avoid non-deterministic randomness in model code. Use fixed random seeds or pure functions to ensure the deterministic pipeline in
_internal/generation.pyproduces consistent hashes.
Implementation Examples
Configure a high-speed SSD cache:
import os
os.environ["CADGEN_CACHE_DIR"] = "/mnt/ssd/cadgen-cache"
# Subsequent cadgen calls will reuse cached artifacts from SSD
Run heavy STEP compilation with automatic parallelization:
import cadgen.step as step
# Assuming my_model.py defines a @step model
step.compile("my_model.py", out="part.step") # Fast on cache hit
# The daemon automatically spawns workers if a cache miss occurs
Tune the daemon for an 8-core workstation:
export CADGEN_DAEMON_WORKERS=8
export CADGEN_DAEMON_TIMEOUT=600 # 10 minutes max per job
cadgen viewer start # Starts the viewer and tuned daemon
Explicitly bypass the cache for debugging:
import os
os.environ["CADGEN_DISABLE_CACHE"] = "1"
import cadgen.step as step
step.compile("my_model.py", out="fresh.step")
Summary
- The content-addressed store in
packages/cadgen/src/cadgen/store/eliminates redundant computation by reusing immutable cached objects stored under~/.cache/cadgen. - Deterministic builds in
_internal/generation.pyensure identical inputs produce byte-identical outputs, maximizing cache hit rates across runs. - The daemon pool in
cadgen/daemon/executors.pyparallelizes heavy kernel operations like STEP compilation across configurable workers. - Environment variables
CADGEN_CACHE_DIRandCADGEN_DAEMON_WORKERSprovide direct control over storage location and parallelism levels. - Avoid
CADGEN_DISABLE_CACHEin production environments, as it forces full recomputation of all geometry and eliminates performance gains.
Frequently Asked Questions
Where does text-to-CAD store its cache by default?
The system stores immutable content-addressed objects in ~/.cache/cadgen by default. You can override this location by setting the CADGEN_CACHE_DIR environment variable to any fast storage path, as demonstrated in tests/python/support/cad_test_roots.py.
How many daemon workers should I configure for optimal performance?
Set CADGEN_DAEMON_WORKERS to match your machine's physical CPU core count. For an 8-core machine, use export CADGEN_DAEMON_WORKERS=8. Monitor actual CPU utilization via cadgen daemon status to verify the pool is saturating available cores without causing system oversubscription.
Why are my builds not hitting the cache despite identical code?
Cache misses typically indicate non-determinism in your build pipeline or excessive churn in decorator side-car metadata. Ensure your model code uses fixed random seeds and that decorator arguments remain stable between runs, as the deterministic pipeline requires byte-identical inputs to generate matching content hashes.
Can I disable caching for debugging purposes?
Yes, set export CADGEN_DISABLE_CACHE=1 to bypass the content-addressed store. This forces the daemon to execute every kernel operation fresh, which is useful for debugging geometry generation issues but dramatically reduces performance. Remove this variable for normal operations to restore caching benefits.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →