How cadgen Supports Parallel Building: Component Process Pools and BREP Serialization
cadgen accelerates CAD model generation by distributing component-level geometry work across a configurable process pool, using BREP payload serialization and environment-variable-driven worker counts to optimize throughput.
The cadgen package from the earthtojake/text-to-cad repository implements sophisticated parallel building capabilities that allow complex CAD models to be constructed significantly faster by isolating independent component workloads. Unlike single-threaded generation, cadgen's architecture delegates geometry computation to stateless worker processes that communicate via deterministic binary payloads. This design ensures thread-safe operations while maximizing CPU utilization through intelligent process pool management.
Component-Level Parallel Architecture
At the core of cadgen's parallel building system is the component package (packages/cadgen/src/cadgen/_internal/component_package.py), which handles the serialization logic and worker allocation strategies necessary for safe multiprocessing.
BREP Payload Serialization for Stateless Workers
cadgen achieves process isolation by serializing each component to a BREP payload that contains pure geometry data without triangulation. This payload acts as a deterministic representation of the component's shape, enabling reconstruction in separate processes.
The system uses two critical functions for this translation: _shape_brep_bytes serializes the shape into bytes, while _build123d_shape_from_brep_bytes reconstructs the geometry object on the worker side. Because the BREP format captures only the mathematical definition of the shape, workers produce identical results regardless of which process performs the computation.
Dynamic Worker Allocation Logic
The parallel_worker_count function (lines 24-42 in component_package.py) determines the optimal number of processes to spawn based on workload size and environment configuration:
def parallel_worker_count(work_count: int, *, env_var: str) -> int:
env_value = os.environ.get(env_var, "").strip()
if env_value:
try:
requested = int(env_value)
except ValueError:
requested = 0
return max(1, min(requested, work_count)) if requested > 1 else 1
if work_count < 6:
return 1
return max(1, min((os.cpu_count() or 2) - 2, work_count, 8))
The helper _component_build_worker_count forwards the missing-component count to this function. When the CADGEN_COMPONENT_WORKERS environment variable is set, its value (clamped between 1 and the pending job count) overrides the default heuristic. Otherwise, cadgen uses a conservative strategy: 1 worker for fewer than 6 jobs, or up to min(cpu_count - 2, 8) for larger workloads.
Worker Pool Implementation
Once the worker count is established, cadgen orchestrates parallel execution using Python's concurrent.futures.ProcessPoolExecutor, with clear separation between the worker logic and the main thread coordination.
Isolated Component Build Workers
Each worker process executes _build_component_surf_worker (lines 98-119 in component_package.py), which receives a tuple of (payload, cid, out_surf, face_colors). The function reconstructs the shape from BREP bytes, optionally restores face-color mappings, and writes the final .surf artifact atomically:
def _build_component_surf_worker(
args: tuple[bytes, str, str, dict | None],
) -> tuple[str, str | None]:
payload, cid, out_surf, face_colors = args
try:
shape = _build123d_shape_from_brep_bytes(payload)
# ... color restoration logic ...
_write_component_artifacts_atomic(shape, Path(out_surf), cad_ref=cid, brep_bytes=payload)
return (cid, None)
except Exception as exc:
return (cid, f"{type(exc).__name__}: {exc}")
If deserialization fails, the worker returns a PAYLOAD_UNREADABLE marker, allowing the parent process to retry the component in-process rather than failing the entire build.
Main Thread Orchestration
The store layer (packages/cadgen/src/cadgen/store/build.py) coordinates the parallel execution. It collects all missing component payloads, queries _component_build_worker_count for the appropriate pool size, and maps the worker function across the payload list:
missing_payloads = [...] # list of (payload, cid, out_surf, face_colors)
workers = _component_build_worker_count(len(missing_payloads))
with ProcessPoolExecutor(max_workers=workers) as pool:
for cid, err in pool.map(_build_component_surf_worker, missing_payloads):
if err:
# Log error or trigger in-process retry
pass
The main thread aggregates (cid, error) tuples, handling any reported failures without blocking on individual component completion.
Daemon-Level Parallelism for Document Compilation
Beyond component-level parallelism, cadgen supports document-level concurrency through its daemon architecture (packages/cadgen/src/cadgen/daemon/pool.py). This separate worker pool handles complete document compilation jobs submitted via submit_compile in daemon/executors.py.
The daemon respects the CADGEN_DAEMON_WORKERS environment variable using the same parallel_worker_count logic, allowing multiple CAD documents to be compiled simultaneously. While component workers focus on geometry reconstruction, daemon workers manage the broader build pipeline, enabling horizontal scaling across independent models.
Configuration and Usage Examples
Control parallel behavior through environment variables when invoking cadgen via CLI or Python API:
# Launch a build with 4 dedicated component workers
CADGEN_COMPONENT_WORKERS=4 cadgen build path/to/model.py
For programmatic usage, set variables before importing the build module:
import os
from cadgen.store.build import build_tree_from_compound
os.environ["CADGEN_COMPONENT_WORKERS"] = "6"
build_tree_from_compound("my_model") # Spawns ProcessPoolExecutor internally
To leverage daemon-level parallelism for multiple documents:
from cadgen.daemon.executors import submit_compile
# Submit independent jobs; daemon runs them in parallel based on CADGEN_DAEMON_WORKERS
submit_compile("model_a.py")
submit_compile("model_b.py")
Summary
- cadgen implements parallel building through isolated process pools that execute component-level geometry work.
- BREP serialization ensures deterministic, stateless worker processes by transmitting pure geometry payloads without triangulation data.
- Dynamic worker allocation defaults to 1 worker for small jobs (<6 components) and scales up to
min(cpu_count - 2, 8)for larger workloads, respecting theCADGEN_COMPONENT_WORKERSenvironment variable. - Error isolation allows unreadable payloads to be flagged for in-process retry without failing the entire build.
- Dual parallelism modes support both component-level generation (via
ProcessPoolExecutorinstore/build.py) and document-level compilation (via the daemon pool indaemon/pool.py).
Frequently Asked Questions
How does cadgen determine the number of parallel workers by default?
cadgen uses the parallel_worker_count function in component_package.py to apply a conservative heuristic: builds with fewer than 6 pending components use a single worker to minimize overhead, while larger workloads scale up to the lesser of (cpu_count - 2) or 8 workers. This default prevents resource exhaustion on systems with many CPU cores while still accelerating substantial builds.
What is the difference between component workers and daemon workers?
Component workers (governed by CADGEN_COMPONENT_WORKERS) are short-lived processes spawned by store/build.py to reconstruct individual geometry components from BREP payloads. Daemon workers (controlled by CADGEN_DAEMON_WORKERS) are persistent processes managed by daemon/pool.py that handle complete document compilation jobs. The former optimizes single-model generation, while the latter enables concurrent processing of multiple independent models.
How does cadgen handle failures in parallel worker processes?
Worker functions return (cid, error) tuples rather than raising exceptions directly. If _build_component_surf_worker encounters a deserialization failure, it returns a PAYLOAD_UNREADABLE marker; for other exceptions, it returns the error string. The main thread in store/build.py inspects these results and can retry failed components in-process, ensuring that transient worker failures do not abort the entire build.
Can parallel building be disabled in cadgen?
Yes. Setting CADGEN_COMPONENT_WORKERS=1 forces single-process execution for component generation, while CADGEN_DAEMON_WORKERS=1 restricts the daemon to sequential document processing. Alternatively, workloads with fewer than 6 components automatically default to 1 worker unless explicitly overridden.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →