What "Resume Is Not Retry" Means in Apache Maka's Architecture
In Apache Maka's distributed mail-processing platform, resuming a job continues execution from a paused state without incrementing attempt counters or triggering backoff logic, whereas retrying creates a new attempt after failure with incremented counters and fresh scheduling delays.
Apache Maka is a high-throughput mail delivery system built around a rigorous state-machine architecture. The principle that "resume is not retry" governs how the platform handles workflow interruptions versus execution failures. This separation ensures accurate reliability metrics, prevents duplicate processing, and maintains idempotent delivery guarantees.
Resume vs. Retry: State Machine Semantics
The distinction centers on how JobStateMachine interprets state transitions. Each operation serves a distinct purpose with different implications for job metadata and execution context.
Resume Continues Paused Execution
A resume operation occurs when a job transitions from PAUSED to RUNNING, typically due to back-pressure relief or administrative intervention. This path:
- Preserves the existing attempt-id and original timestamps stored in Job.java
- Retrieves the execution context from CheckpointStore without re-initialization
- Maintains partially processed data and in-memory state
- Does not invoke RetryPolicy or reset failure counters
Retry Handles Failure Recovery
A retry operation triggers when a job enters the FAILED state and the system classifies the error as transient. Unlike resume, this represents a new execution attempt:
- Increments the attempt counter and generates a new attempt-id
- Applies exponential backoff calculations via RetryPolicy
- May trigger rollback of partial work before re-execution
- Creates distinct audit entries and metrics separate from previous attempts
Architectural Implementation
Apache Maka enforces this separation through three specialized components that isolate pause handling from failure recovery.
JobStateMachine Governs Transitions
The JobStateMachine.java file defines valid state transitions and their side effects. When processing a state change, the machine determines whether to preserve or increment attempt metadata based on the target state.
// Logic pattern from JobStateMachine.java
public void transition(Job job, JobState target) {
switch (target) {
case RESUMED:
// Preserve original attempt context
job.setAttemptId(job.getAttemptId());
processingEngine.process(job);
break;
case RETRY:
// Increment for new attempt
job.setAttemptId(job.getAttemptId() + 1);
retryPolicy.apply(job);
processingEngine.process(job);
break;
// ... other states
}
}
ResumeHandler Restores Context
The ResumeHandler.java component manages restoration of paused jobs by interfacing with the persistence layer. It reloads execution state without touching retry logic.
// Implementation from ResumeHandler.java
public void handleResume(Job job) {
Checkpoint cp = checkpointStore.load(job.getId());
job.restoreFromCheckpoint(cp);
// Forward without modifying retry counters
stateMachine.transition(job, JobState.RESUMED);
}
Because ResumeHandler never consults RetryPolicy, the operation remains strictly a continuation rather than a new attempt.
RetryPolicy Isolates Failure Logic
The RetryPolicy.java encapsulates backoff strategies and maximum attempt enforcement. This component is only invoked during FAILED state transitions, ensuring temporary pauses never trigger exponential delays or attempt exhaustion.
// Scheduling logic from RetryPolicy.java
public void apply(Job job) {
long delay = computeBackoff(job.getAttemptId());
scheduler.schedule(
() -> stateMachine.transition(job, JobState.RETRY),
delay
);
}
Operational Impact
Separating resume from retry prevents metric pollution in monitoring systems. If resuming incremented attempt counters, dashboards would incorrectly report administrative pauses as delivery failures. This architecture also guarantees idempotent processing—a resumed job picks up exactly where checkpointing occurred without re-executing committed work.
Summary
- Resume preserves the original attempt-id and loads context from CheckpointStore, treating the operation as a simple state continuation.
- Retry increments attempt counters and applies RetryPolicy backoff calculations only after explicit failure state transitions.
- JobStateMachine.java enforces distinct code paths that prevent conflation of paused and failed states.
- This separation eliminates duplicate work risks and ensures reliability metrics accurately reflect actual delivery attempts versus capacity management pauses.
Frequently Asked Questions
Does resuming a job count against the maximum retry limit?
No. ResumeHandler does not interact with RetryPolicy or modify the attempt counter in Job.java. Only transitions through the FAILED state consume retry attempts, ensuring temporary pauses for back-pressure or maintenance do not exhaust a job's permitted failure budget.
How does Apache Maka ensure data consistency during a resume?
The platform uses CheckpointStore.java to persist job context, processing offsets, and partial results when a job enters the PAUSED state. When ResumeHandler processes the continuation, it restores this checkpointed state exactly, allowing the job to proceed without reprocessing already-committed data.
Can an administrator convert a paused job into a retry?
While possible through manual state manipulation, the architecture discourages this by design. Forcing a paused job through the FAILED state to trigger RetryPolicy would increment the attempt counter and potentially trigger unnecessary rollback logic, distorting operational metrics and wasting processing resources.
Why does the distinction matter for observability?
Mixing resume and retry operations would conflate two different operational signals: system health issues requiring investigation (retries) versus capacity management requiring simple continuation (pauses). By keeping these paths separate in JobStateMachine, Apache Maka ensures alerting systems can accurately distinguish between transient resource constraints and actual delivery failures.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →