How the Open Agents Workflow SDK Enables Durable Multi-Step Execution
The Open Agents Workflow SDK transforms standard async functions into durable orchestrators that survive server crashes, automatically retry failed operations, and resume from external events using a lightweight "use workflow" directive and persistent state management.
The vercel-labs/open-agents repository leverages Vercel's Workflow DevKit (the workflow NPM package, version ^4.2.0-beta.72 as declared in apps/web/package.json) to build resilient AI agent processes. Unlike standard Node.js functions that lose state on deployment or crashes, the Open Agents workflow SDK enables durable multi-step execution by persisting function state after every await and providing built-in retry logic for individual steps.
Core Concepts of Durable Execution
The "use workflow" Directive
A function whose first line is "use workflow" executes inside a sandboxed VM whose state is persisted on every await. This directive, implemented in files like apps/web/app/workflows/chat.ts, signals the runtime to treat the function as a durable orchestrator rather than a standard serverless handler.
When the runtime encounters "use workflow", it wraps the function in a persistent execution context. If the server crashes, deploys new code, or restarts, the workflow runtime can restart the function and restore its exact state—including all local variables and execution position—from the last persisted snapshot.
Step Functions with "use step"
Inside a workflow, you define step functions whose first line is "use step". Unlike the workflow itself, steps run in a regular Node.js environment with full access to the filesystem, npm modules, and external APIs. The apps/web/app/workflows/chat.ts file demonstrates this pattern by breaking complex AI operations into discrete, retry-able units.
Steps automatically receive retry logic, caching, and streaming support from the SDK. When a step fails, the workflow doesn't crash; instead, the runtime applies exponential backoff retry policies before either succeeding or escalating the error based on its classification.
Starting and Managing Workflow Runs
Initializing Runs with start()
The start function from workflow/api creates a new workflow execution and returns a unique runId. The implementation in apps/web/lib/sandbox/lifecycle-kick.ts demonstrates how Open Agents initiates workflows:
import { start } from "workflow/api";
// Creates a new run and returns runId
const runId = await start({
workflow: "chat",
params: { message: "Hello" }
});
The runId serves as the durable identifier for this specific execution, allowing the system to query status, resume after pauses, or stream results even if the initiating process disconnects.
Database Persistence
The SDK relies on a database schema defined in apps/web/lib/db/schema.ts to store execution state. Two critical tables enable durability:
workflowRuns: Stores the top-level run metadata, including current status, workflow type, and serialization checkpointsworkflowRunSteps: Persists individual step outputs, retry counts, and execution timestamps
This persistence layer ensures that workflow state survives server restarts, deployments, or infrastructure failures. When a workflow resumes, the runtime reconstructs the execution context from these database records rather than starting from scratch.
Resilience and Error Handling
Automatic Retry Mechanisms
The Open Agents workflow SDK implements intelligent retry logic at the step level. When a step function throws an error, the runtime automatically attempts to re-execute the step with exponential backoff before propagating the failure to the parent workflow.
This mechanism is particularly valuable for AI operations that may encounter transient API rate limits or network interruptions. The retry configuration respects the "use step" directive's boundary, ensuring that idempotent operations can safely repeat while the workflow maintains its overall position.
FatalError vs RetryableError
The SDK provides specific error classes to control retry behavior, documented in .agents/skills/workflow/SKILL.md:
FatalError: Immediately aborts the entire workflow run permanently, preventing any further retries or continuationRetryableError: Signals the runtime to retry the current step with backoff, preserving the workflow's durable state
import { FatalError, RetryableError } from "workflow";
// Permanent failure - stops the workflow
throw new FatalError("Invalid API key");
// Transient failure - will retry
throw new RetryableError("Rate limit exceeded");
This distinction allows developers to model business logic failures (fatal) separately from infrastructure hiccups (retryable), optimizing resource usage and debugging clarity.
Advanced Capabilities
Streaming Results with getWritable
Long-running workflows can stream partial results to clients using getWritable from the workflow package. The apps/web/app/workflows/chat.ts implementation demonstrates this pattern for AI chat responses:
import { getWritable } from "workflow";
const writable = getWritable();
const writer = writable.getWriter();
// Stream tokens as they arrive from the AI
writer.write(JSON.stringify({ token: "Hello" }));
writer.write(JSON.stringify({ token: " world" }));
writer.close();
This capability enables real-time user experiences while the workflow continues executing subsequent steps. The writable stream persists across server restarts, ensuring clients don't lose connection state during infrastructure changes.
External Event Resumption
Workflows can pause execution to wait for external triggers using resumeHook and resumeWebhook. These functions, documented in .agents/skills/workflow/SKILL.md, allow the workflow to suspend without consuming resources while awaiting user approval, third-party webhooks, or scheduled timers.
When the external event fires, the runtime uses the stored runId to locate the paused workflow and resume execution at the exact await point. This mechanism effectively turns long-polling or webhook-based integrations into durable, serverless-friendly operations.
Timers and Fetch
The SDK provides sandbox-compatible replacements for standard JavaScript primitives:
sleep: ReplacessetTimeoutinside workflows, allowing the VM to pause without blocking the event loopfetch: A sandbox-compatible fetch implementation that can attach toglobalThis.fetchwhen needed
import { sleep, fetch } from "workflow";
// Pause for 5 seconds (durable across restarts)
await sleep("5s");
// Sandbox-safe HTTP request
globalThis.fetch = fetch;
const response = await fetch("https://api.example.com/data");
These utilities ensure that workflows remain deterministic and serializable while interacting with external systems.
Real-World Implementation
The Chat Workflow Example
The apps/web/app/workflows/chat.ts file demonstrates production-grade durable execution. This workflow orchestrates a multi-turn AI conversation using the SDK's core primitives:
- Entry point: Declares
"use workflow"to enable durability - Step boundaries: Isolates AI model calls and database writes into
"use step"functions - Streaming: Uses
getWritableto send tokens to the UI in real-time - Persistence: Stores run state via the
workflowRunstable for recovery
This implementation showcases how complex, long-running AI agent interactions become fault-tolerant without manual state management or complex retry logic.
Summary
-
"use workflow" transforms async functions into durable orchestrators that survive server crashes and deployments by persisting state on every
await. -
"use step" isolates individual operations with automatic retry, caching, and streaming, running in standard Node.js while maintaining workflow durability.
-
Database persistence via
workflowRunsandworkflowRunStepstables ensures state recovery across infrastructure failures. -
Advanced primitives like
getWritable,resumeHook,sleep, andfetchenable streaming, external event waiting, and sandbox-compatible I/O without breaking durability. -
Error classification through
FatalErrorandRetryableErrorallows precise control over retry behavior and workflow termination.
Frequently Asked Questions
What is the difference between "use workflow" and "use step"?
"use workflow" marks the main orchestration function that runs in a sandboxed VM with automatic state persistence, while "use step" marks individual task functions that execute in standard Node.js with full system access but automatic retry and caching. The workflow survives crashes and restarts; steps are the retryable units of work that can access the filesystem and npm modules.
How does the Open Agents workflow handle server restarts?
The SDK persists workflow state to the database (workflowRuns and workflowRunSteps tables defined in apps/web/lib/db/schema.ts) after every await boundary. When a server restarts, the runtime reads the last checkpoint from the database and resumes execution at the exact line where the workflow paused, reconstructing all local variables and the call stack without losing progress.
Can workflows stream data to clients while continuing to process?
Yes. The getWritable function from the workflow package returns a WritableStream that persists across server restarts. As demonstrated in apps/web/app/workflows/chat.ts, steps can write partial results (like AI tokens) to the client while subsequent steps continue processing, enabling real-time chat interfaces with durable backend execution.
What happens when a step encounters an error?
The runtime automatically retries steps that throw RetryableError with exponential backoff, preserving the workflow's durable state between attempts. If a step throws FatalError or exhausts retry limits, the entire workflow run aborts permanently. This distinction allows transient failures (API rate limits) to resolve automatically while permanent failures (invalid credentials) terminate the process immediately.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →