# How Orca Handles Multi-Threading: Worker Threads, Child Processes, and Native Parallelism

> Discover how Orca handles multi-threading using worker threads, child processes, and native parallelism to maintain a responsive UI during intensive tasks like speech recognition and I/O. Learn more.

- Repository: [Stably/orca](https://github.com/stablyai/orca)
- Tags: internals
- Published: 2026-05-25

---

**Orca isolates CPU-intensive speech recognition and terminal I/O into separate threads and processes, keeping the Electron UI responsive through a combination of Node.js `worker_threads`, forked child processes, and native OS-level parallelism.**

The `stablyai/orca` repository implements a strict separation of concerns for concurrency, ensuring that heavy computational workloads never block the main application thread. By distributing tasks across distinct execution contexts—from JavaScript worker threads to native inference engines—Orca maintains a fluid user experience even during resource-intensive operations like real-time speech-to-text processing.

## Node.js Worker Threads for Speech Recognition

### Main Thread Orchestration

In [`src/main/speech/stt-service.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-service.ts), the main Electron thread instantiates a dedicated `Worker` that executes [`src/main/speech/stt-worker.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-worker.ts) in isolation. This pattern prevents the UI from freezing during the expensive initialization of the `sherpa-onnx` native addon. The service manages the worker's lifecycle through message passing, using `worker.on('message')` to receive transcription results while the worker uses `parentPort.postMessage` to stream data back to the main process.

### Worker Thread Message Protocol

The worker thread operates on a strict command-based protocol defined in [`src/main/speech/stt-worker.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-worker.ts):

```typescript
// src/main/speech/stt-worker.ts
parentPort?.on('message', (msg) => {
  switch (msg.type) {
    case 'init':   handleInit(msg);   break;
    case 'feed':   handleFeed(msg);   break;
    case 'stop':   handleStop();      break;
    case 'teardown': handleTeardown(); break;
  }
});

```

This architecture allows the main thread to feed audio chunks asynchronously while the worker handles blocking recognition logic on a separate thread.

## Forked Child Processes for PTY Handling

For pseudo-terminal operations, Orca uses `child_process.fork()` rather than worker threads to isolate blocking system calls. Files including **[`src/relay/pty-shell-launch.ts`](https://github.com/stablyai/orca/blob/main/src/relay/pty-shell-launch.ts)**, **[`src/relay/pty-handler.ts`](https://github.com/stablyai/orca/blob/main/src/relay/pty-handler.ts)**, and **[`src/relay/pty-shell-utils.ts`](https://github.com/stablyai/orca/blob/main/src/relay/pty-shell-utils.ts)** spawn separate Node processes that own the PTY:

```typescript
// src/relay/pty-shell-launch.ts
import { fork } from 'child_process';
const ptyDaemon = fork(join(__dirname, 'pty-daemon.js'), [], { stdio: 'pipe' });
ptyDaemon.on('message', (msg) => { /* IPC handling */ });

```

The forked process executes low-level terminal I/O while the main process communicates via IPC messages, ensuring that shell operations never stall the UI event loop.

## Native OS Threads for Model Inference

The `sherpa-onnx` native module creates its own OS threads for model inference, independent of JavaScript's concurrency model. Orca configures the thread pool size by passing `numThreads` to the recognizer configuration (lines 31-33 in [`src/main/speech/stt-worker.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-worker.ts)). The JavaScript layer merely forwards this parameter; the actual parallel matrix operations execute on native threads without consuming JavaScript execution resources.

## Automatic Resource Cleanup via Idle Teardown

To prevent long-running sessions from pinning large ONNX models in memory, `SttService.scheduleIdleTeardown()` implements an idle timeout mechanism using the `IDLE_WORKER_TEARDOWN_MS` constant. After dictation ends, this timer counts down before sending a `teardown` command to the worker. The worker then releases the native recognizer instance, allowing the V8 thread to be garbage-collected while keeping the model warm for repeated dictations during active use.

## Summary

- **Worker Threads**: [`src/main/speech/stt-service.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-service.ts) creates a `Worker` from [`src/main/speech/stt-worker.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-worker.ts) to run speech recognition off the main thread, communicating via `postMessage` and `on('message')` events.
- **Child Processes**: PTY operations in [`src/relay/pty-shell-launch.ts`](https://github.com/stablyai/orca/blob/main/src/relay/pty-shell-launch.ts) and [`src/relay/pty-handler.ts`](https://github.com/stablyai/orca/blob/main/src/relay/pty-handler.ts) use `fork()` to isolate blocking terminal I/O from the Electron main process.
- **Native Threading**: The `sherpa-onnx` addon manages its own OS threads for inference, configured via `numThreads` in lines 31-33 of the worker configuration.
- **Lifecycle Management**: The `SttService` class employs `scheduleIdleTeardown()` with `IDLE_WORKER_TEARDOWN_MS` to release native resources after periods of inactivity, preventing memory leaks.

## Frequently Asked Questions

### Why does Orca use worker_threads for speech recognition instead of running it on the main thread?

Running the `sherpa-onnx` speech engine on the main thread would block the Electron UI during model loading and audio inference. By delegating this work to a `Worker` in [`src/main/speech/stt-worker.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-worker.ts), Orca maintains responsive user interactions while the CPU-intensive recognition runs in parallel on a separate thread.

### What is the difference between forked child processes and worker threads in Orca's architecture?

Orca uses **worker threads** for shared-memory tasks like audio processing where message passing is sufficient, while **forked child processes** ([`src/relay/pty-shell-launch.ts`](https://github.com/stablyai/orca/blob/main/src/relay/pty-shell-launch.ts)) handle PTY operations that require process isolation and blocking system calls. Child processes provide stronger isolation for system-level I/O, whereas worker threads are lighter for CPU-bound JavaScript tasks.

### How does Orca prevent memory leaks from long-running speech models?

The `SttService` class implements `scheduleIdleTeardown()`, which starts an `IDLE_WORKER_TEARDOWN_MS` timer after dictation ends. When the timer fires, the service sends a `teardown` message to the worker in [`src/main/speech/stt-worker.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-worker.ts), explicitly releasing the native `sherpa-onnx` recognizer and allowing garbage collection of the V8 thread.

### Where is the thread count configured for the native speech recognition engine?

The `numThreads` parameter is set in the recognizer configuration within [`src/main/speech/stt-worker.ts`](https://github.com/stablyai/orca/blob/main/src/main/speech/stt-worker.ts) at lines 31-33. This value controls how many OS threads the `sherpa-onnx` native module creates for parallel inference, independent of the single Node.js worker thread that hosts the addon.