Paperclip AI Agent Work Verification with Diffs and Screenshots: A Deep Source‑Code Analysis

Paperclip verifies AI agent work by persisting unified diffs of code changes alongside Playwright‑generated screenshots, storing both artefacts in the run model for side‑by‑side review in the UI.

The Paperclip repository implements a rigorous verification system for AI‑driven development workflows. When an autonomous agent edits files, every change is captured as a textual diff and every UI state is frozen as a screenshot. This dual‑artefact approach ensures that human reviewers can audit what the agent changed and how those changes appear—without trusting the agent's own reports.

How Diffs Are Generated and Stored

When an agent issues an edit tool call, the server records line‑level changes (additions, removals, or context lines) as discrete diff entries. These entries are assembled into a TaskChatDiff object and attached to the running tool record.

The diff structure appears in ui/src/components/task-chat/task-chat-fixtures.ts:

{
  kind: "diff",
  ts: "2026-04-06T12:00:02.000Z",
  changeType: "add",
  text: "+function formatDiffBlock(lines: string[]) {"
}

Server‑side persistence happens in server/src/services/smoke-lab.ts. The service attaches the artefact reference to the run model under screenshotArtifactRef (a shared field reused for screenshots):

screenshotArtifactRef: input.screenshotArtifactRef ?? null,

The diff payload travels with the tool record, while the artefact reference enables later retrieval from storage.

Diff Rendering in the UI

The client‑side renderer constructs fenced diff blocks for display. In ui/src/lib/issue-chat-messages.ts, the system builds markdown‑style diff fences from the stored entries. The TaskChat UI model defined in ui/src/components/task-chat/task-chat-model.ts includes the optional diff field:

/** Optional unified‑diff‑ish lines for display only. */
diff?: TaskChatDiff;

How Screenshots Are Captured

Paperclip's end‑to‑end test suite drives the agent's actions through Playwright, capturing full‑page PNG screenshots after each significant interaction.

The screenshot() helper in tests/e2e/smoke-lab.spec.ts handles path generation and capture:

async function screenshot(page: Page, scenario: SmokeLabScenario, step: string) {
  const path = `${SCREENSHOT_DIR}/${scenario.id}-${step}.png`;
  await page.screenshot({ path, fullPage: true });
  return path;
}

Screenshot Artefact Validation

Every successful step must include a screenshot reference. The test suite enforces this invariant:

expect(steps.every((step) => step.screenshotArtifactRef?.kind === "playwright_screenshot")).toBe(true);

The reference object has a discriminated structure:

{ kind: "playwright_screenshot", path: ".../scenario-id-step.png" }

Visual Regression Safeguards

Beyond runtime verification, Paperclip maintains baseline screenshots through a dedicated visual snapshot suite in tests/storybook-visual/storybook-visual.spec.ts. This suite captures one screenshot per Storybook component per theme.

The Playwright configuration in tests/storybook-visual/playwright.config.ts ensures screenshots are taken only after components fully settle, eliminating timing‑related flakiness and providing deterministic regression detection.

Complete Verification Pipeline

The Paperclip AI agent work verification system operates across five integrated stages:

  1. Agent execution — The agent emits tool calls, including edit operations with line‑level changes.
  2. Diff assembly — The server transforms individual changes into TaskChatDiff entries attached to tool records.
  3. Screenshot capture — Playwright executes the run, calls screenshot() after each step, and returns the PNG path.
  4. Artefact persistence — The server stores screenshotArtifactRef alongside diffs in the run model.
  5. UI rendering — The client fetches run data, displays diff blocks as syntax‑highlighted code, and renders screenshot thumbnails with click‑to‑enlarge behaviour.

Code Examples in Context

These patterns appear throughout the codebase. Here is how the pieces connect in practice:

Capture a screenshot during an e2e test:

const screenshotPath = await screenshot(page, scenario, "step-1");

Attach both diff and screenshot to the server payload:

await fetch("/api/runs", {
  method: "POST",
  body: JSON.stringify({
    tool: "edit",
    diff: [{ kind: "add", text: "+const foo = 42;" }],
    screenshotArtifactRef: {
      kind: "playwright_screenshot",
      path: screenshotPath,
    },
  }),
});

Render a diff block in a React component:

function DiffBlock({ diff }: { diff: TaskChatDiff }) {
  return <pre className="diff">{diff.lines.map(l => l.text).join("\n")}</pre>;
}

Summary

  • Unified diffs capture every line‑level change from agent edit tool calls, stored as TaskChatDiff entries.
  • Playwright screenshots provide visual proof of UI state after each step, persisted via screenshotArtifactRef objects.
  • Dual verification combines textual and visual artefacts in the server/src/services/smoke-lab.ts run model.
  • Deterministic baselines from tests/storybook-visual/ enable long‑term regression tracking.
  • Integrated UI rendering in ui/src/components/task-chat/ presents both artefacts for human review.

Frequently Asked Questions

What triggers diff generation in Paperclip?

Diff generation is triggered by edit tool calls from the AI agent. Each call emits line‑level change records (add, remove, context) that the server assembles into a TaskChatDiff object according to ui/src/components/task-chat/task-chat-model.ts.

Why does Paperclip use both diffs and screenshots instead of one or the other?

Textual diffs precisely answer what code changed, while screenshots answer how the UI looks after those changes. Together they eliminate ambiguity: a diff might pass while breaking layout, or a screenshot might hide logic errors visible only in code. The smoke-lab.ts service stores both artefact types for complete accountability.

How does Paperclip prevent flaky screenshots in visual regression tests?

The tests/storybook-visual/playwright.config.ts configures Playwright to capture screenshots only after components reach a settled state. Combined with one‑shot‑per‑theme discipline in storybook-visual.spec.ts, this produces deterministic baselines unsusceptible to animation or loading timing variations.

Can screenshot artefacts be retrieved independently of the run transcript?

Yes. The screenshotArtifactRef stored in the run model contains a kind discriminator and absolute path, allowing direct storage access. The UI uses this reference to render thumbnails and full‑size images without parsing the entire transcript.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →