How the OpenMAIC Job Store Manages Render Job Lifecycle and Cancellation
The OpenMAIC render service maintains a dedicated Job Store layer that tracks every render request from creation to completion, enabling robust lifecycle management and immediate cancellation through AbortSignal propagation independent of the HTTP layer.
The OpenMAIC repository implements a sophisticated render architecture that separates job persistence from request handling. By isolating render job lifecycle management into a dedicated Job Store, the system ensures reliable state tracking, per-user concurrency limits, and safe cancellation semantics even when clients disconnect unexpectedly. This design allows the RenderCoordinator and RenderExecutor to operate autonomously while maintaining observable, cancellable job states.
Core Responsibilities of the Job Store
The InMemoryJobStore class in render-service/src/job-store.ts provides the foundational persistence layer for all render operations. It maintains an internal Map structure that stores RenderJobRecord instances and exposes methods that drive the entire lifecycle.
Create, Read, Update, and Delete Operations
The Job Store implements standard CRUD operations with specific responsibilities:
-
InMemoryJobStore.create– Initializes a newRenderJobRecordwith metadata including ID, user ID, timestamps, and project folder location. This method is invoked atjob-store.ts#L42-L45when the HTTP handler processes aPOST /renderrequest. -
InMemoryJobStore.get– Retrieves the current record by ID, returning the stored object ornullif not found. The poll endpoint uses this atjob-store.ts#L46-L48to return status updates to clients. -
InMemoryJobStore.update– Merges partial patches into existing records and refreshes theupdatedAtMstimestamp. This handles status transitions, progress updates, and error attachments atjob-store.ts#L50-L54. -
InMemoryJobStore.remove– Deletes the map entry when a job is explicitly purged, implemented atjob-store.ts#L56-L58.
Concurrency Guard and Automatic Cleanup
Beyond basic persistence, the Job Store enforces resource limits and automated maintenance:
Per-user concurrency counting is handled by countActiveForUser, which traverses the internal Map and counts jobs where !isTerminal (neither succeeded, failed, nor cancelled). This prevents individual users from overwhelming the system with too many simultaneous renders, as seen at job-store.ts#L64-L69.
TTL-based cleanup runs via the sweep() method, executed every minute through setInterval. This method identifies terminal jobs older than ttlMs, removes them from the Map, and invokes the optional onReap callback to delete temporary project directories. The sweeper implementation resides at job-store.ts#L72-L80.
Job Cancellation Architecture
Cancellation in OpenMAIC uses the standard AbortController/AbortSignal pattern, allowing jobs to terminate gracefully at multiple stages of the render job lifecycle.
AbortSignal Integration
Every render execution receives an AbortSignal through the RenderExecutionRequest interface defined at types.ts#L91-L93:
export interface RenderExecutionRequest {
// ... other fields ...
/** User or coordinator cancellation. */
signal: AbortSignal;
// ... other fields ...
}
The RenderExecutor.execute method checks this signal immediately upon entry. If request.signal.aborted is true, the executor returns a cancelled result without launching any worker processes, as implemented at render-executor.ts#L88-L94.
Coordinator-Level Cancellation Logic
The RenderCoordinator.cancel(id) method at render-coordinator.ts#L29-L55 handles external cancellation requests:
- Controller lookup – Retrieves the
AbortControllerassociated with the job ID fromthis.controllers.get(id). - Signal propagation – Calls
controller.abort(), which triggers the executor's abort handling logic. - Queue removal – If the job remains queued, it is removed from the queue, the per-user concurrency guard is decremented, and the job record is immediately patched to cancelled status with a
RenderCancelledFailureobject. - Resource cleanup – Invokes the cleanup callback to remove the temporary project directory.
Lifecycle Finalization
Regardless of termination reason—success, failure, or cancellation—the coordinator calls finishEvent at render-coordinator.ts#L71-L78. This method logs the outcome, removes timing entries, and ensures idempotent closure of the job lifecycle to prevent double-counting or duplicate cleanup operations.
After cancellation, the Job Store record reflects the terminal state defined in RenderJobStatus at types.ts#L14-L16:
{
status: 'cancelled',
currentStage: 'cancelled',
failure: { code: 'cancelled', message: 'Render cancelled' }
}
End-to-End Render Job Lifecycle Walkthrough
The complete lifecycle of a render job progresses through distinct phases orchestrated by the Job Store and coordinator:
- Creation – Client sends
POST /render, triggeringJobStore.createto persist aqueuedrecord. - Enqueuing – The coordinator stores an
AbortControllerfor the job and places it in the processing queue. - Polling – Client requests
GET /render/:id, which callsJobStore.getto return current status and progress. - Execution – When scheduled,
RenderExecutor.executeruns with the controller's signal, transitioning status torunning. - Cancellation paths:
- User-initiated:
POST /render/:id/canceltriggerscoordinator.cancel(id), aborting the controller. - Signal-initiated: External timeouts or navigation events abort the signal before execution begins, causing immediate return of a cancelled result.
- User-initiated:
- Completion – Upon resolution (cancelled, succeeded, or failed), the coordinator updates the Job Store via
JobStore.updateand callsfinishEventfor logging and cleanup. - Reaping – The TTL sweeper eventually removes terminal records after the configured delay, invoking
onReapto delete temporary files.
Implementation Examples
Creating a Render Job
The HTTP handler uses the Job Store to persist new requests:
// Simplified snippet from the POST /render handler
const jobId = crypto.randomUUID();
await jobStore.create({
id: jobId,
userId,
status: 'queued',
progress: 0,
currentStage: 'queued',
createdAtMs: Date.now(),
updatedAtMs: Date.now(),
projectDir,
});
coordinator.enqueue(jobId, /* record */);
Source: job-store.ts#L42-L45
Polling Job Status
Status retrieval for client updates:
// GET /render/:id
const record = await jobStore.get(jobId);
if (!record) return response.notFound();
return response.json({
status: record.status,
progress: record.progress,
stage: record.currentStage,
error: record.error,
});
Source: job-store.ts#L46-L48
Cancelling a Job
External cancellation endpoint:
// POST /render/:id/cancel
const cancelled = await coordinator.cancel(jobId);
if (!cancelled) return response.notFound();
return response.json({ cancelled: true });
Source: render-coordinator.ts#L29-L55
Executor Short-Circuit on Abort
Early exit when cancellation occurs before or during execution:
async execute(request: RenderExecutionRequest) {
if (request.signal.aborted) {
return {
status: 'cancelled',
failure: { code: 'cancelled', message: 'Render cancelled' },
};
}
// ... normal render logic ...
}
Source: render-executor.ts#L88-L94
Summary
- The Job Store abstraction in
job-store.tsprovides atomic CRUD operations, per-user concurrency counting, and TTL-based automatic cleanup through thesweep()method. - Cancellation safety is achieved via standard
AbortControllerpatterns, allowing immediate termination whether a job is queued or actively rendering. - Resource hygiene is guaranteed through dual cleanup paths: immediate directory removal upon cancellation in
RenderCoordinator.cancel, and deferred reaping by the TTL sweeper'sonReapcallback. - State consistency is maintained through
RenderCoordinator.finishEvent, ensuring idempotent lifecycle closure and accurate final status persistence in the Job Store.
Frequently Asked Questions
How does the Job Store enforce per-user concurrency limits?
The countActiveForUser method iterates through the internal Map and counts jobs where the status is not terminal (not succeeded, failed, or cancelled). This count is checked before new jobs are enqueued to prevent any single user from exceeding their allocated render slots.
What happens to temporary project files when a render job is cancelled?
When RenderCoordinator.cancel is invoked, it immediately triggers a cleanup callback that deletes the temporary project directory associated with the job. Additionally, if cleanup is deferred, the TTL sweeper's onReap callback removes files when the job record is eventually purged from memory.
Can a render job be cancelled after execution has already started?
Yes. Each running job maintains an AbortController stored by the coordinator. Calling cancel(id) aborts this controller, and the executor checks signal.aborted at strategic points—including the start of execution—to exit early and return a cancelled status before completing the render.
How does the system prevent memory leaks from accumulated job records?
The sweep() method runs every 60 seconds via setInterval, scanning for terminal jobs exceeding the configured ttlMs duration. These records are deleted from the internal Map, and the optional onReap hook ensures associated temporary resources are freed, preventing unbounded memory growth in long-running service instances.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →