# Paperclip AI Agent Work Verification with Diffs and Screenshots: A Deep Source‑Code Analysis

> Verify Paperclip AI agent work with diffs and screenshots. See code changes and UI artifacts side-by-side for deep source-code analysis. Enhance your AI agent development.

- Repository: [Paperclip/paperclip](https://github.com/paperclipai/paperclip)
- Tags: deep-dive
- Published: 2026-08-12

---

**Paperclip verifies AI agent work by persisting unified diffs of code changes alongside Playwright‑generated screenshots, storing both artefacts in the run model for side‑by‑side review in the UI.**

The [Paperclip](https://github.com/paperclipai/paperclip) repository implements a rigorous verification system for AI‑driven development workflows. When an autonomous agent edits files, every change is captured as a textual diff and every UI state is frozen as a screenshot. This dual‑artefact approach ensures that human reviewers can audit *what* the agent changed and *how* those changes appear—without trusting the agent's own reports.

## How Diffs Are Generated and Stored

When an agent issues an **`edit`** tool call, the server records line‑level changes (additions, removals, or context lines) as discrete diff entries. These entries are assembled into a `TaskChatDiff` object and attached to the running tool record.

The diff structure appears in [`ui/src/components/task-chat/task-chat-fixtures.ts`](https://github.com/paperclipai/paperclip/blob/main/ui/src/components/task-chat/task-chat-fixtures.ts):

```typescript
{
  kind: "diff",
  ts: "2026-04-06T12:00:02.000Z",
  changeType: "add",
  text: "+function formatDiffBlock(lines: string[]) {"
}

```

Server‑side persistence happens in [`server/src/services/smoke-lab.ts`](https://github.com/paperclipai/paperclip/blob/main/server/src/services/smoke-lab.ts). The service attaches the artefact reference to the run model under `screenshotArtifactRef` (a shared field reused for screenshots):

```typescript
screenshotArtifactRef: input.screenshotArtifactRef ?? null,

```

The diff payload travels with the tool record, while the artefact reference enables later retrieval from storage.

### Diff Rendering in the UI

The client‑side renderer constructs fenced `diff` blocks for display. In [`ui/src/lib/issue-chat-messages.ts`](https://github.com/paperclipai/paperclip/blob/main/ui/src/lib/issue-chat-messages.ts), the system builds markdown‑style diff fences from the stored entries. The `TaskChat` UI model defined in [`ui/src/components/task-chat/task-chat-model.ts`](https://github.com/paperclipai/paperclip/blob/main/ui/src/components/task-chat/task-chat-model.ts) includes the optional diff field:

```typescript
/** Optional unified‑diff‑ish lines for display only. */
diff?: TaskChatDiff;

```

## How Screenshots Are Captured

Paperclip's end‑to‑end test suite drives the agent's actions through **Playwright**, capturing full‑page PNG screenshots after each significant interaction.

The `screenshot()` helper in [`tests/e2e/smoke-lab.spec.ts`](https://github.com/paperclipai/paperclip/blob/main/tests/e2e/smoke-lab.spec.ts) handles path generation and capture:

```typescript
async function screenshot(page: Page, scenario: SmokeLabScenario, step: string) {
  const path = `${SCREENSHOT_DIR}/${scenario.id}-${step}.png`;
  await page.screenshot({ path, fullPage: true });
  return path;
}

```

### Screenshot Artefact Validation

Every successful step must include a screenshot reference. The test suite enforces this invariant:

```typescript
expect(steps.every((step) => step.screenshotArtifactRef?.kind === "playwright_screenshot")).toBe(true);

```

The reference object has a discriminated structure:

```typescript
{ kind: "playwright_screenshot", path: ".../scenario-id-step.png" }

```

## Visual Regression Safeguards

Beyond runtime verification, Paperclip maintains baseline screenshots through a dedicated visual snapshot suite in [`tests/storybook-visual/storybook-visual.spec.ts`](https://github.com/paperclipai/paperclip/blob/main/tests/storybook-visual/storybook-visual.spec.ts). This suite captures one screenshot per Storybook component per theme.

The Playwright configuration in [`tests/storybook-visual/playwright.config.ts`](https://github.com/paperclipai/paperclip/blob/main/tests/storybook-visual/playwright.config.ts) ensures screenshots are taken only after components fully settle, eliminating timing‑related flakiness and providing deterministic regression detection.

## Complete Verification Pipeline

The Paperclip AI agent work verification system operates across five integrated stages:

1. **Agent execution** — The agent emits tool calls, including `edit` operations with line‑level changes.
2. **Diff assembly** — The server transforms individual changes into `TaskChatDiff` entries attached to tool records.
3. **Screenshot capture** — Playwright executes the run, calls `screenshot()` after each step, and returns the PNG path.
4. **Artefact persistence** — The server stores `screenshotArtifactRef` alongside diffs in the run model.
5. **UI rendering** — The client fetches run data, displays diff blocks as syntax‑highlighted code, and renders screenshot thumbnails with click‑to‑enlarge behaviour.

## Code Examples in Context

These patterns appear throughout the codebase. Here is how the pieces connect in practice:

**Capture a screenshot during an e2e test:**

```typescript
const screenshotPath = await screenshot(page, scenario, "step-1");

```

**Attach both diff and screenshot to the server payload:**

```typescript
await fetch("/api/runs", {
  method: "POST",
  body: JSON.stringify({
    tool: "edit",
    diff: [{ kind: "add", text: "+const foo = 42;" }],
    screenshotArtifactRef: {
      kind: "playwright_screenshot",
      path: screenshotPath,
    },
  }),
});

```

**Render a diff block in a React component:**

```typescript
function DiffBlock({ diff }: { diff: TaskChatDiff }) {
  return <pre className="diff">{diff.lines.map(l => l.text).join("\n")}</pre>;
}

```

## Summary

- **Unified diffs** capture every line‑level change from agent `edit` tool calls, stored as `TaskChatDiff` entries.
- **Playwright screenshots** provide visual proof of UI state after each step, persisted via `screenshotArtifactRef` objects.
- **Dual verification** combines textual and visual artefacts in the [`server/src/services/smoke-lab.ts`](https://github.com/paperclipai/paperclip/blob/main/server/src/services/smoke-lab.ts) run model.
- **Deterministic baselines** from `tests/storybook-visual/` enable long‑term regression tracking.
- **Integrated UI rendering** in `ui/src/components/task-chat/` presents both artefacts for human review.

## Frequently Asked Questions

### What triggers diff generation in Paperclip?

Diff generation is triggered by **`edit`** tool calls from the AI agent. Each call emits line‑level change records (add, remove, context) that the server assembles into a `TaskChatDiff` object according to [`ui/src/components/task-chat/task-chat-model.ts`](https://github.com/paperclipai/paperclip/blob/main/ui/src/components/task-chat/task-chat-model.ts).

### Why does Paperclip use both diffs and screenshots instead of one or the other?

Textual diffs precisely answer *what code changed*, while screenshots answer *how the UI looks after those changes*. Together they eliminate ambiguity: a diff might pass while breaking layout, or a screenshot might hide logic errors visible only in code. The [`smoke-lab.ts`](https://github.com/paperclipai/paperclip/blob/main/smoke-lab.ts) service stores both artefact types for complete accountability.

### How does Paperclip prevent flaky screenshots in visual regression tests?

The [`tests/storybook-visual/playwright.config.ts`](https://github.com/paperclipai/paperclip/blob/main/tests/storybook-visual/playwright.config.ts) configures Playwright to capture screenshots only after components reach a settled state. Combined with one‑shot‑per‑theme discipline in [`storybook-visual.spec.ts`](https://github.com/paperclipai/paperclip/blob/main/storybook-visual.spec.ts), this produces deterministic baselines unsusceptible to animation or loading timing variations.

### Can screenshot artefacts be retrieved independently of the run transcript?

Yes. The `screenshotArtifactRef` stored in the run model contains a `kind` discriminator and absolute `path`, allowing direct storage access. The UI uses this reference to render thumbnails and full‑size images without parsing the entire transcript.