# How cadgen Supports Parallel Building: Component Process Pools and BREP Serialization

> Discover how cadgen accelerates CAD generation through parallel building. Learn about component process pools and BREP serialization for optimized throughput.

- Repository: [earthtojake/text-to-cad](https://github.com/earthtojake/text-to-cad)
- Tags: internals
- Published: 2026-09-11

---

**cadgen accelerates CAD model generation by distributing component-level geometry work across a configurable process pool, using BREP payload serialization and environment-variable-driven worker counts to optimize throughput.**

The `cadgen` package from the earthtojake/text-to-cad repository implements sophisticated parallel building capabilities that allow complex CAD models to be constructed significantly faster by isolating independent component workloads. Unlike single-threaded generation, cadgen's architecture delegates geometry computation to stateless worker processes that communicate via deterministic binary payloads. This design ensures thread-safe operations while maximizing CPU utilization through intelligent process pool management.

## Component-Level Parallel Architecture

At the core of cadgen's parallel building system is the **component package** ([`packages/cadgen/src/cadgen/_internal/component_package.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadgen/src/cadgen/_internal/component_package.py)), which handles the serialization logic and worker allocation strategies necessary for safe multiprocessing.

### BREP Payload Serialization for Stateless Workers

cadgen achieves process isolation by serializing each component to a **BREP payload** that contains pure geometry data without triangulation. This payload acts as a deterministic representation of the component's shape, enabling reconstruction in separate processes.

The system uses two critical functions for this translation: `_shape_brep_bytes` serializes the shape into bytes, while `_build123d_shape_from_brep_bytes` reconstructs the geometry object on the worker side. Because the BREP format captures only the mathematical definition of the shape, workers produce identical results regardless of which process performs the computation.

### Dynamic Worker Allocation Logic

The `parallel_worker_count` function (lines 24-42 in [`component_package.py`](https://github.com/earthtojake/text-to-cad/blob/main/component_package.py)) determines the optimal number of processes to spawn based on workload size and environment configuration:

```python
def parallel_worker_count(work_count: int, *, env_var: str) -> int:
    env_value = os.environ.get(env_var, "").strip()
    if env_value:
        try:
            requested = int(env_value)
        except ValueError:
            requested = 0
        return max(1, min(requested, work_count)) if requested > 1 else 1
    if work_count < 6:
        return 1
    return max(1, min((os.cpu_count() or 2) - 2, work_count, 8))

```

The helper `_component_build_worker_count` forwards the missing-component count to this function. When the `CADGEN_COMPONENT_WORKERS` environment variable is set, its value (clamped between 1 and the pending job count) overrides the default heuristic. Otherwise, cadgen uses a conservative strategy: **1 worker for fewer than 6 jobs**, or up to `min(cpu_count - 2, 8)` for larger workloads.

## Worker Pool Implementation

Once the worker count is established, cadgen orchestrates parallel execution using Python's `concurrent.futures.ProcessPoolExecutor`, with clear separation between the worker logic and the main thread coordination.

### Isolated Component Build Workers

Each worker process executes `_build_component_surf_worker` (lines 98-119 in [`component_package.py`](https://github.com/earthtojake/text-to-cad/blob/main/component_package.py)), which receives a tuple of `(payload, cid, out_surf, face_colors)`. The function reconstructs the shape from BREP bytes, optionally restores face-color mappings, and writes the final `.surf` artifact atomically:

```python
def _build_component_surf_worker(
    args: tuple[bytes, str, str, dict | None],
) -> tuple[str, str | None]:
    payload, cid, out_surf, face_colors = args
    try:
        shape = _build123d_shape_from_brep_bytes(payload)
        # ... color restoration logic ...

        _write_component_artifacts_atomic(shape, Path(out_surf), cad_ref=cid, brep_bytes=payload)
        return (cid, None)
    except Exception as exc:
        return (cid, f"{type(exc).__name__}: {exc}")

```

If deserialization fails, the worker returns a `PAYLOAD_UNREADABLE` marker, allowing the parent process to retry the component in-process rather than failing the entire build.

### Main Thread Orchestration

The store layer ([`packages/cadgen/src/cadgen/store/build.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadgen/src/cadgen/store/build.py)) coordinates the parallel execution. It collects all missing component payloads, queries `_component_build_worker_count` for the appropriate pool size, and maps the worker function across the payload list:

```python
missing_payloads = [...]  # list of (payload, cid, out_surf, face_colors)

workers = _component_build_worker_count(len(missing_payloads))
with ProcessPoolExecutor(max_workers=workers) as pool:
    for cid, err in pool.map(_build_component_surf_worker, missing_payloads):
        if err:
            # Log error or trigger in-process retry

            pass

```

The main thread aggregates `(cid, error)` tuples, handling any reported failures without blocking on individual component completion.

## Daemon-Level Parallelism for Document Compilation

Beyond component-level parallelism, cadgen supports **document-level concurrency** through its daemon architecture ([`packages/cadgen/src/cadgen/daemon/pool.py`](https://github.com/earthtojake/text-to-cad/blob/main/packages/cadgen/src/cadgen/daemon/pool.py)). This separate worker pool handles complete document compilation jobs submitted via `submit_compile` in [`daemon/executors.py`](https://github.com/earthtojake/text-to-cad/blob/main/daemon/executors.py).

The daemon respects the `CADGEN_DAEMON_WORKERS` environment variable using the same `parallel_worker_count` logic, allowing multiple CAD documents to be compiled simultaneously. While component workers focus on geometry reconstruction, daemon workers manage the broader build pipeline, enabling horizontal scaling across independent models.

## Configuration and Usage Examples

Control parallel behavior through environment variables when invoking cadgen via CLI or Python API:

```bash

# Launch a build with 4 dedicated component workers

CADGEN_COMPONENT_WORKERS=4 cadgen build path/to/model.py

```

For programmatic usage, set variables before importing the build module:

```python
import os
from cadgen.store.build import build_tree_from_compound

os.environ["CADGEN_COMPONENT_WORKERS"] = "6"
build_tree_from_compound("my_model")  # Spawns ProcessPoolExecutor internally

```

To leverage daemon-level parallelism for multiple documents:

```python
from cadgen.daemon.executors import submit_compile

# Submit independent jobs; daemon runs them in parallel based on CADGEN_DAEMON_WORKERS

submit_compile("model_a.py")
submit_compile("model_b.py")

```

## Summary

- **cadgen** implements parallel building through isolated process pools that execute component-level geometry work.
- **BREP serialization** ensures deterministic, stateless worker processes by transmitting pure geometry payloads without triangulation data.
- **Dynamic worker allocation** defaults to 1 worker for small jobs (<6 components) and scales up to `min(cpu_count - 2, 8)` for larger workloads, respecting the `CADGEN_COMPONENT_WORKERS` environment variable.
- **Error isolation** allows unreadable payloads to be flagged for in-process retry without failing the entire build.
- **Dual parallelism modes** support both component-level generation (via `ProcessPoolExecutor` in [`store/build.py`](https://github.com/earthtojake/text-to-cad/blob/main/store/build.py)) and document-level compilation (via the daemon pool in [`daemon/pool.py`](https://github.com/earthtojake/text-to-cad/blob/main/daemon/pool.py)).

## Frequently Asked Questions

### How does cadgen determine the number of parallel workers by default?

cadgen uses the `parallel_worker_count` function in [`component_package.py`](https://github.com/earthtojake/text-to-cad/blob/main/component_package.py) to apply a conservative heuristic: builds with fewer than 6 pending components use a single worker to minimize overhead, while larger workloads scale up to the lesser of `(cpu_count - 2)` or 8 workers. This default prevents resource exhaustion on systems with many CPU cores while still accelerating substantial builds.

### What is the difference between component workers and daemon workers?

**Component workers** (governed by `CADGEN_COMPONENT_WORKERS`) are short-lived processes spawned by [`store/build.py`](https://github.com/earthtojake/text-to-cad/blob/main/store/build.py) to reconstruct individual geometry components from BREP payloads. **Daemon workers** (controlled by `CADGEN_DAEMON_WORKERS`) are persistent processes managed by [`daemon/pool.py`](https://github.com/earthtojake/text-to-cad/blob/main/daemon/pool.py) that handle complete document compilation jobs. The former optimizes single-model generation, while the latter enables concurrent processing of multiple independent models.

### How does cadgen handle failures in parallel worker processes?

Worker functions return `(cid, error)` tuples rather than raising exceptions directly. If `_build_component_surf_worker` encounters a deserialization failure, it returns a `PAYLOAD_UNREADABLE` marker; for other exceptions, it returns the error string. The main thread in [`store/build.py`](https://github.com/earthtojake/text-to-cad/blob/main/store/build.py) inspects these results and can retry failed components in-process, ensuring that transient worker failures do not abort the entire build.

### Can parallel building be disabled in cadgen?

Yes. Setting `CADGEN_COMPONENT_WORKERS=1` forces single-process execution for component generation, while `CADGEN_DAEMON_WORKERS=1` restricts the daemon to sequential document processing. Alternatively, workloads with fewer than 6 components automatically default to 1 worker unless explicitly overridden.