Paperclip AI Agent Work Verification with Diffs and Screenshots: A Deep Source‑Code Analysis
Paperclip verifies AI agent work by persisting unified diffs of code changes alongside Playwright‑generated screenshots, storing both artefacts in the run model for side‑by‑side review in the UI.
The Paperclip repository implements a rigorous verification system for AI‑driven development workflows. When an autonomous agent edits files, every change is captured as a textual diff and every UI state is frozen as a screenshot. This dual‑artefact approach ensures that human reviewers can audit what the agent changed and how those changes appear—without trusting the agent's own reports.
How Diffs Are Generated and Stored
When an agent issues an edit tool call, the server records line‑level changes (additions, removals, or context lines) as discrete diff entries. These entries are assembled into a TaskChatDiff object and attached to the running tool record.
The diff structure appears in ui/src/components/task-chat/task-chat-fixtures.ts:
{
kind: "diff",
ts: "2026-04-06T12:00:02.000Z",
changeType: "add",
text: "+function formatDiffBlock(lines: string[]) {"
}
Server‑side persistence happens in server/src/services/smoke-lab.ts. The service attaches the artefact reference to the run model under screenshotArtifactRef (a shared field reused for screenshots):
screenshotArtifactRef: input.screenshotArtifactRef ?? null,
The diff payload travels with the tool record, while the artefact reference enables later retrieval from storage.
Diff Rendering in the UI
The client‑side renderer constructs fenced diff blocks for display. In ui/src/lib/issue-chat-messages.ts, the system builds markdown‑style diff fences from the stored entries. The TaskChat UI model defined in ui/src/components/task-chat/task-chat-model.ts includes the optional diff field:
/** Optional unified‑diff‑ish lines for display only. */
diff?: TaskChatDiff;
How Screenshots Are Captured
Paperclip's end‑to‑end test suite drives the agent's actions through Playwright, capturing full‑page PNG screenshots after each significant interaction.
The screenshot() helper in tests/e2e/smoke-lab.spec.ts handles path generation and capture:
async function screenshot(page: Page, scenario: SmokeLabScenario, step: string) {
const path = `${SCREENSHOT_DIR}/${scenario.id}-${step}.png`;
await page.screenshot({ path, fullPage: true });
return path;
}
Screenshot Artefact Validation
Every successful step must include a screenshot reference. The test suite enforces this invariant:
expect(steps.every((step) => step.screenshotArtifactRef?.kind === "playwright_screenshot")).toBe(true);
The reference object has a discriminated structure:
{ kind: "playwright_screenshot", path: ".../scenario-id-step.png" }
Visual Regression Safeguards
Beyond runtime verification, Paperclip maintains baseline screenshots through a dedicated visual snapshot suite in tests/storybook-visual/storybook-visual.spec.ts. This suite captures one screenshot per Storybook component per theme.
The Playwright configuration in tests/storybook-visual/playwright.config.ts ensures screenshots are taken only after components fully settle, eliminating timing‑related flakiness and providing deterministic regression detection.
Complete Verification Pipeline
The Paperclip AI agent work verification system operates across five integrated stages:
- Agent execution — The agent emits tool calls, including
editoperations with line‑level changes. - Diff assembly — The server transforms individual changes into
TaskChatDiffentries attached to tool records. - Screenshot capture — Playwright executes the run, calls
screenshot()after each step, and returns the PNG path. - Artefact persistence — The server stores
screenshotArtifactRefalongside diffs in the run model. - UI rendering — The client fetches run data, displays diff blocks as syntax‑highlighted code, and renders screenshot thumbnails with click‑to‑enlarge behaviour.
Code Examples in Context
These patterns appear throughout the codebase. Here is how the pieces connect in practice:
Capture a screenshot during an e2e test:
const screenshotPath = await screenshot(page, scenario, "step-1");
Attach both diff and screenshot to the server payload:
await fetch("/api/runs", {
method: "POST",
body: JSON.stringify({
tool: "edit",
diff: [{ kind: "add", text: "+const foo = 42;" }],
screenshotArtifactRef: {
kind: "playwright_screenshot",
path: screenshotPath,
},
}),
});
Render a diff block in a React component:
function DiffBlock({ diff }: { diff: TaskChatDiff }) {
return <pre className="diff">{diff.lines.map(l => l.text).join("\n")}</pre>;
}
Summary
- Unified diffs capture every line‑level change from agent
edittool calls, stored asTaskChatDiffentries. - Playwright screenshots provide visual proof of UI state after each step, persisted via
screenshotArtifactRefobjects. - Dual verification combines textual and visual artefacts in the
server/src/services/smoke-lab.tsrun model. - Deterministic baselines from
tests/storybook-visual/enable long‑term regression tracking. - Integrated UI rendering in
ui/src/components/task-chat/presents both artefacts for human review.
Frequently Asked Questions
What triggers diff generation in Paperclip?
Diff generation is triggered by edit tool calls from the AI agent. Each call emits line‑level change records (add, remove, context) that the server assembles into a TaskChatDiff object according to ui/src/components/task-chat/task-chat-model.ts.
Why does Paperclip use both diffs and screenshots instead of one or the other?
Textual diffs precisely answer what code changed, while screenshots answer how the UI looks after those changes. Together they eliminate ambiguity: a diff might pass while breaking layout, or a screenshot might hide logic errors visible only in code. The smoke-lab.ts service stores both artefact types for complete accountability.
How does Paperclip prevent flaky screenshots in visual regression tests?
The tests/storybook-visual/playwright.config.ts configures Playwright to capture screenshots only after components reach a settled state. Combined with one‑shot‑per‑theme discipline in storybook-visual.spec.ts, this produces deterministic baselines unsusceptible to animation or loading timing variations.
Can screenshot artefacts be retrieved independently of the run transcript?
Yes. The screenshotArtifactRef stored in the run model contains a kind discriminator and absolute path, allowing direct storage access. The UI uses this reference to render thumbnails and full‑size images without parsing the entire transcript.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →