How the OpenMAIC Job Store Manages Render Job Lifecycle and Cancellation

The OpenMAIC render service maintains a dedicated Job Store layer that tracks every render request from creation to completion, enabling robust lifecycle management and immediate cancellation through AbortSignal propagation independent of the HTTP layer.

The OpenMAIC repository implements a sophisticated render architecture that separates job persistence from request handling. By isolating render job lifecycle management into a dedicated Job Store, the system ensures reliable state tracking, per-user concurrency limits, and safe cancellation semantics even when clients disconnect unexpectedly. This design allows the RenderCoordinator and RenderExecutor to operate autonomously while maintaining observable, cancellable job states.

Core Responsibilities of the Job Store

The InMemoryJobStore class in render-service/src/job-store.ts provides the foundational persistence layer for all render operations. It maintains an internal Map structure that stores RenderJobRecord instances and exposes methods that drive the entire lifecycle.

Create, Read, Update, and Delete Operations

The Job Store implements standard CRUD operations with specific responsibilities:

  • InMemoryJobStore.create – Initializes a new RenderJobRecord with metadata including ID, user ID, timestamps, and project folder location. This method is invoked at job-store.ts#L42-L45 when the HTTP handler processes a POST /render request.

  • InMemoryJobStore.get – Retrieves the current record by ID, returning the stored object or null if not found. The poll endpoint uses this at job-store.ts#L46-L48 to return status updates to clients.

  • InMemoryJobStore.update – Merges partial patches into existing records and refreshes the updatedAtMs timestamp. This handles status transitions, progress updates, and error attachments at job-store.ts#L50-L54.

  • InMemoryJobStore.remove – Deletes the map entry when a job is explicitly purged, implemented at job-store.ts#L56-L58.

Concurrency Guard and Automatic Cleanup

Beyond basic persistence, the Job Store enforces resource limits and automated maintenance:

Per-user concurrency counting is handled by countActiveForUser, which traverses the internal Map and counts jobs where !isTerminal (neither succeeded, failed, nor cancelled). This prevents individual users from overwhelming the system with too many simultaneous renders, as seen at job-store.ts#L64-L69.

TTL-based cleanup runs via the sweep() method, executed every minute through setInterval. This method identifies terminal jobs older than ttlMs, removes them from the Map, and invokes the optional onReap callback to delete temporary project directories. The sweeper implementation resides at job-store.ts#L72-L80.

Job Cancellation Architecture

Cancellation in OpenMAIC uses the standard AbortController/AbortSignal pattern, allowing jobs to terminate gracefully at multiple stages of the render job lifecycle.

AbortSignal Integration

Every render execution receives an AbortSignal through the RenderExecutionRequest interface defined at types.ts#L91-L93:

export interface RenderExecutionRequest {
  // ... other fields ...
  /** User or coordinator cancellation. */
  signal: AbortSignal;
  // ... other fields ...
}

The RenderExecutor.execute method checks this signal immediately upon entry. If request.signal.aborted is true, the executor returns a cancelled result without launching any worker processes, as implemented at render-executor.ts#L88-L94.

Coordinator-Level Cancellation Logic

The RenderCoordinator.cancel(id) method at render-coordinator.ts#L29-L55 handles external cancellation requests:

  1. Controller lookup – Retrieves the AbortController associated with the job ID from this.controllers.get(id).
  2. Signal propagation – Calls controller.abort(), which triggers the executor's abort handling logic.
  3. Queue removal – If the job remains queued, it is removed from the queue, the per-user concurrency guard is decremented, and the job record is immediately patched to cancelled status with a RenderCancelledFailure object.
  4. Resource cleanup – Invokes the cleanup callback to remove the temporary project directory.

Lifecycle Finalization

Regardless of termination reason—success, failure, or cancellation—the coordinator calls finishEvent at render-coordinator.ts#L71-L78. This method logs the outcome, removes timing entries, and ensures idempotent closure of the job lifecycle to prevent double-counting or duplicate cleanup operations.

After cancellation, the Job Store record reflects the terminal state defined in RenderJobStatus at types.ts#L14-L16:

{
  status: 'cancelled',
  currentStage: 'cancelled',
  failure: { code: 'cancelled', message: 'Render cancelled' }
}

End-to-End Render Job Lifecycle Walkthrough

The complete lifecycle of a render job progresses through distinct phases orchestrated by the Job Store and coordinator:

  1. Creation – Client sends POST /render, triggering JobStore.create to persist a queued record.
  2. Enqueuing – The coordinator stores an AbortController for the job and places it in the processing queue.
  3. Polling – Client requests GET /render/:id, which calls JobStore.get to return current status and progress.
  4. Execution – When scheduled, RenderExecutor.execute runs with the controller's signal, transitioning status to running.
  5. Cancellation paths:
    • User-initiated: POST /render/:id/cancel triggers coordinator.cancel(id), aborting the controller.
    • Signal-initiated: External timeouts or navigation events abort the signal before execution begins, causing immediate return of a cancelled result.
  6. Completion – Upon resolution (cancelled, succeeded, or failed), the coordinator updates the Job Store via JobStore.update and calls finishEvent for logging and cleanup.
  7. Reaping – The TTL sweeper eventually removes terminal records after the configured delay, invoking onReap to delete temporary files.

Implementation Examples

Creating a Render Job

The HTTP handler uses the Job Store to persist new requests:

// Simplified snippet from the POST /render handler
const jobId = crypto.randomUUID();
await jobStore.create({
  id: jobId,
  userId,
  status: 'queued',
  progress: 0,
  currentStage: 'queued',
  createdAtMs: Date.now(),
  updatedAtMs: Date.now(),
  projectDir,
});
coordinator.enqueue(jobId, /* record */);

Source: job-store.ts#L42-L45

Polling Job Status

Status retrieval for client updates:

// GET /render/:id
const record = await jobStore.get(jobId);
if (!record) return response.notFound();
return response.json({
  status: record.status,
  progress: record.progress,
  stage: record.currentStage,
  error: record.error,
});

Source: job-store.ts#L46-L48

Cancelling a Job

External cancellation endpoint:

// POST /render/:id/cancel
const cancelled = await coordinator.cancel(jobId);
if (!cancelled) return response.notFound();
return response.json({ cancelled: true });

Source: render-coordinator.ts#L29-L55

Executor Short-Circuit on Abort

Early exit when cancellation occurs before or during execution:

async execute(request: RenderExecutionRequest) {
  if (request.signal.aborted) {
    return {
      status: 'cancelled',
      failure: { code: 'cancelled', message: 'Render cancelled' },
    };
  }
  // ... normal render logic ...
}

Source: render-executor.ts#L88-L94

Summary

  • The Job Store abstraction in job-store.ts provides atomic CRUD operations, per-user concurrency counting, and TTL-based automatic cleanup through the sweep() method.
  • Cancellation safety is achieved via standard AbortController patterns, allowing immediate termination whether a job is queued or actively rendering.
  • Resource hygiene is guaranteed through dual cleanup paths: immediate directory removal upon cancellation in RenderCoordinator.cancel, and deferred reaping by the TTL sweeper's onReap callback.
  • State consistency is maintained through RenderCoordinator.finishEvent, ensuring idempotent lifecycle closure and accurate final status persistence in the Job Store.

Frequently Asked Questions

How does the Job Store enforce per-user concurrency limits?

The countActiveForUser method iterates through the internal Map and counts jobs where the status is not terminal (not succeeded, failed, or cancelled). This count is checked before new jobs are enqueued to prevent any single user from exceeding their allocated render slots.

What happens to temporary project files when a render job is cancelled?

When RenderCoordinator.cancel is invoked, it immediately triggers a cleanup callback that deletes the temporary project directory associated with the job. Additionally, if cleanup is deferred, the TTL sweeper's onReap callback removes files when the job record is eventually purged from memory.

Can a render job be cancelled after execution has already started?

Yes. Each running job maintains an AbortController stored by the coordinator. Calling cancel(id) aborts this controller, and the executor checks signal.aborted at strategic points—including the start of execution—to exit early and return a cancelled status before completing the render.

How does the system prevent memory leaks from accumulated job records?

The sweep() method runs every 60 seconds via setInterval, scanning for terminal jobs exceeding the configured ttlMs duration. These records are deleted from the internal Map, and the optional onReap hook ensures associated temporary resources are freed, preventing unbounded memory growth in long-running service instances.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →