How the Native SDK Automation Server Enables AI Agent Testing and UI Snapshots

The Native SDK automation server enables AI agent testing by exposing a drop-box file protocol that lets external scripts queue commands and read back deterministic, text-based UI snapshots of every window, widget, and runtime metric.

The vercel-labs/native repository implements this server in Zig to bridge external test harnesses with a live Native app. By treating the filesystem as a lock-free message bus, the Native SDK automation server isolates the agent from the app process while still enabling full UI inspection and control. This architecture makes it safe for automated CI pipelines, large-scale AI-driven exploration, and snapshot-based regression testing.

Drop-Box Directory and Protocol Handshake

All communication happens through a shared directory defined in src/automation/protocol.zig. The constant default_dir pins the drop-box to .zig-cache/native-sdk-automation, which both the CLI client and the running app watch.

pub const default_dir = ".zig-cache/native-sdk-automation"; // protocol.zig L3

Before accepting traffic, the server advertises its protocol version in src/automation/server.zig. When the runtime boots, it writes a ready=true line that embeds the current version—7—so the CLI can abort early on mismatch.

try writer.print("ready=true protocol={d} …\n", .{protocol.version}); // server.zig L40-L46

This handshake guarantees that the AI agent and the app speak the same contract before any commands or snapshots are exchanged.

You can initialise the server inside the native app with a single call:

const Server = @import("src/automation/server.zig").Server;

var server = Server.init(std.testing.io,
    ".zig-cache/native-sdk-automation", // directory
    "native-sdk");                       // title shown in logs

Publishing UI Snapshots and Test Artifacts

When the runtime requests a publish cycle, the server writes three artifacts into the drop-box. Each write uses a std.Io.Writer.Allocating instance with pre-allocated capacity—16 KB for snapshots and 1 KB for window metadata—followed by a single atomic writePath call.

  • snapshot.txt — a full UI dump produced by snapshot.writeText starting at line 72 of src/automation/snapshot.zig.
  • accessibility.txt — an assistive-technology tree produced by snapshot.writeA11yText starting at line 2 of src/automation/snapshot.zig.
  • bridge-response.txt / provenance.txt — one-off responses handled by publishBridgeResponse and publishProvenanceResponse in server.zig lines 61-78.
var writer = try std.Io.Writer.Allocating.initCapacity(...);
try snapshot.writeText(input_value, &writer.writer);
try writePath(self.io, self.path("snapshot.txt", &path_buffer), writer.written());

A typical publish call looks like this:

var win = Window{
    .title = "Demo",
    .bounds = geometry.RectF.init(0, 0, 800, 600),
};
var widget = snapshot.Widget{
    .window_id = 1,
    .view_label = "main",
    .id = 100,
    .role = "button",
    .name = "ClickMe",
    .bounds = geometry.RectF.init(100, 100, 120, 40),
    .focused = true,
    .enabled = true,
};

try server.publish(.{
    .windows = &[_]snapshot.Window{win},
    .widgets = &[_]snapshot.Widget{widget},
});

For visual regression, the agent can issue a screenshot command. The runtime then calls publishScreenshot in server.zig (lines 80-97), which writes a temporary PNG and renames it atomically so no poller ever reads a truncated image.

Command Queue: FIFO Dispatch and Crash Recovery

The server consumes work from a flat queue of command-<seq>.txt files dropped by external agents. Sequence numbers are generated by protocol.queueFileName in protocol.zig lines 57-60.

const name = try protocol.queueFileName(sequence, &name_buffer);

Dispatch is strictly FIFO. The helper queueHeadName in server.zig (lines 65-84) scans the directory for the smallest sequence number and returns the oldest entry.

fn queueHeadName(self: Server, name_buffer: []u8) ?[]const u8 { … } // server.zig L65-L84

The runtime calls takeCommand once per frame. This routine reads the head entry, deletes it immediately—the ACK—and returns a parsed protocol.Command. If the file is still being written and lacks a trailing newline, takeCommand defers processing and re-queues the watcher.

pub fn takeCommand(self: Server, buffer: []u8) !?protocol.Command { … } // server.zig L111-L142

Inside the app loop, consumption looks like this:

var read_buf: [256]u8 = undefined;
if (try server.takeCommand(&read_buf)) |command| {
    // Dispatch based on command.action …
    switch (command.action) {
        .widget_click => |*c| std.debug.print("Clicked widget: {s}\n", .{c.value}),
        else => {},
    }
}

To handle crashed writers, the server runs reapAbandonedEntry in server.zig (lines 86-95). Any incomplete file older than abandoned_entry_reap_ns is removed automatically. The protocol also caps total queued commands at 8 via max_queued_commands in protocol.zig lines 49-50, forcing writers to retry rather than allowing unbounded growth.

pub const max_queued_commands: usize = 8; // protocol.zig L49-L50

AI Agent Testing Loop: Drive, Wait, and Observe

An AI agent interacts with the app through a simple four-phase loop. This decouples the agent from the process memory, making the Native SDK automation server ideal for untrusted or large-scale AI agents.

Injecting Commands

The agent composes a command line with protocol.commandLine and writes it to a new command-<seq>.txt file.

var cmd_buf: [128]u8 = undefined;
const cmd = try protocol.commandLine("widget-click", "main 100", &cmd_buf);
try writePath(std.testing.io,
    server.path("command-1.txt", &path_buf), cmd);

Waiting for Runtime Processing

After dropping the file, the agent polls hasPendingCommand in server.zig (lines 44-48) or watches the directory. The entry disappears only after the runtime acknowledges it.

while (server.hasPendingCommand()) std.time.sleep(50_000_000); // 50 ms

Reading Snapshots for Verification

Once the command is consumed, the agent reads snapshot.txt or accessibility.txt to verify state, collect widget IDs, or feed the hierarchy into its own model.

var snap_buf: [64*1024]u8 = undefined;
const snap_text = try readPath(std.testing.io,
    server.path("snapshot.txt", &path_buf), &snap_buf);
// `snap_text` now contains lines like:
//   window @w1 "Demo" …
//   widget @w1/main#100 role=button name="ClickMe" …

Capturing Screenshots

When the agent needs pixel-level verification, it queues a screenshot command. The resulting PNG is delivered atomically by publishScreenshot, guaranteeing that external readers never see a half-written image.

pub fn publishScreenshot(self: Server, view_label: []const u8, png_bytes: []const u8) !void { … } // server.zig L80-L97

Summary

  • The Native SDK automation server in vercel-labs/native uses a filesystem drop-box at .zig-cache/native-sdk-automation to decouple AI agents from the app process.
  • Protocol version 7 enforces compatibility before any commands are exchanged.
  • The server publishes deterministic, text-based snapshots via snapshot.writeText and snapshot.writeA11yText, plus atomic screenshots via publishScreenshot.
  • A FIFO command queue capped at 8 entries and protected by reapAbandonedEntry ensures safe, recoverable dispatch even when writers crash.
  • AI agents drive the UI by writing command-<seq>.txt files, waiting on hasPendingCommand, and reading back structured UI state for verification or model training.

Frequently Asked Questions

How does the Native SDK automation server guarantee snapshot consistency?

The server never writes snapshots in place. Instead, it uses a std.Io.Writer.Allocating buffer to assemble the full text, then commits the file with a single atomic writePath call. Screenshot binaries follow the same pattern: a temp file is written and renamed only when complete. This ensures that AI agents always read valid, complete artifacts.

What happens if an AI agent crashes while writing a command?

Incomplete command files are tracked by age. If a file sits unfinished longer than the abandoned_entry_reap_ns timeout, the server's reapAbandonedEntry routine—implemented in src/automation/server.zig lines 86-95—deletes it automatically. The queue also hard-limits outstanding commands to max_queued_commands = 8, so a crashed agent cannot exhaust server resources.

Can multiple AI agents queue commands at the same time?

Yes, because the queue relies on unique sequence numbers generated by protocol.queueFileName in src/automation/protocol.zig. However, the runtime processes only one command per frame via takeCommand, and the oldest entry wins. If the queue hits its capacity of 8, additional writers must retry, which naturally serializes load.

How does an AI agent know when a command has been executed?

The agent polls hasPendingCommand or watches the drop-box directory. A command is considered acknowledged only when takeCommand deletes its command-<seq>.txt file. Once the file disappears, the agent can safely read snapshot.txt or accessibility.txt to observe the updated UI state.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →