How to Extract Structured Output from Agent Responses Using `Output.object()`
Use Output.object({ tag, schema }) to declare an XML-style tag that the LLM will wrap around JSON output, which sandcastle then extracts, unwraps from Markdown fences if present, validates against a Standard Schema, and returns as a strongly-typed object.
The sandcastle library provides a type-safe mechanism for extracting structured data from LLM agent stdout without fragile string parsing. By combining XML-style tags with Standard Schema validation, you can enforce contracts on agent responses while maintaining clean separation between conversational text and machine-readable data.
Declaring the Expected Shape with Output.object()
The extraction workflow begins by declaring the expected data structure through the Output namespace, defined in src/Output.ts. The object helper creates a branded definition that carries both the tag name and a Standard Schema validator.
At runtime, the declaration resolves to a simple object { _tag: "object", tag, schema } that run() forwards through the pipeline. At the type level, it establishes a contract that TypeScript uses to infer the return type of the extraction.
import { Output } from "@ai-hero/sandcastle";
import { z } from "zod";
const outputConfig = Output.object({
tag: "result",
schema: z.object({ answer: z.number() }), // Standard-Schema-compatible
});
This implementation resides in [src/Output.ts lines 49‑60](https://github.com/mattpocock/sandcastle/blob/main/src/Output.ts#L49-L60).
Running the Sandbox with Structured Output Constraints
When you invoke run() with an output configuration, the function validates two critical pre-conditions before attempting extraction.
First, maxIterations must be set to 1. The library enforces single-iteration runs only when expecting structured output, as the extraction happens after the iteration completes. This check occurs in [src/run.ts lines 60‑66](https://github.com/mattpocock/sandcastle/blob/main/src/run.ts#L60-L66).
Second, the resolved prompt must contain the opening tag literal <${tag}>. If the prompt lacks this marker, the run throws early. See the validation logic in [src/run.ts lines 96‑100](https://github.com/mattpocock/sandcastle/blob/main/src/run.ts#L96-L100).
Once these checks pass, run() executes the agent, aggregates the stdout, and hands it to extractStructuredOutput() (imported from src/extractStructuredOutput.ts). The parsed value is then attached to RunResult.output.
The Extraction Pipeline: From stdout to Typed Data
The extractStructuredOutput() function performs the heavy lifting of locating, cleaning, and validating the structured payload.
Tag Discovery with findLastTagContent()
The extraction routine scans the combined stdout for the last occurrence of the specified XML-style tag. It captures the content between <tag> and </tag>, returning the inner string or undefined if not found. This logic prevents issues when the LLM generates multiple instances of the tag, ensuring the final answer takes precedence over intermediate thoughts.
Implementation details are in [src/extractStructuredOutput.ts lines 12‑34](https://github.com/mattpocock/sandcastle/blob/main/src/extractStructuredOutput.ts#L12-L34).
Markdown Fence Handling
LLMs frequently wrap JSON output in Markdown code fences (```json ... ``` or plain ``` ... ```). The unwrapFences() helper detects these wrappers and strips them before parsing, ensuring the downstream JSON parser receives clean text.
See the fence detection logic in [src/extractStructuredOutput.ts lines 50‑58](https://github.com/mattpocock/sandcastle/blob/main/src/extractStructuredOutput.ts#L50-L58).
Validation and Error Handling
After cleaning, the text undergoes JSON.parse. Any syntax error triggers a StructuredOutputError. If parsing succeeds, the resulting object passes to the Standard Schema validator supplied in the Output.object declaration.
Validation failures also raise StructuredOutputError, which carries:
- The expected tag name
- The raw matched text for debugging
- The original cause (parse error or validation issues)
- Run context (commits, branch, worktree path)
The core validation logic resides in [src/extractStructuredOutput.ts lines 57‑80](https://github.com/mattpocock/sandcastle/blob/main/src/extractStructuredOutput.ts#L57-L80), while the error type definition appears in [src/Output.ts lines 76‑96](https://github.com/mattpocock/sandcastle/blob/main/src/Output.ts#L76-L96).
Complete Working Example
The following example demonstrates calculating a sum with an agent that returns the result inside a <calcResult> tag:
import { run, Output, StructuredOutputError } from "@ai-hero/sandcastle";
import { z } from "zod";
const prompt = `
Calculate the sum of 12 and 30.
Respond exactly with:
<calcResult>
\`\`\`json
{ "sum": 42 }
\`\`\`
</calcResult>
`;
async function example() {
try {
const { output } = await run({
name: "calc-agent",
sandbox: { tag: "docker" },
prompt,
maxIterations: 1, // Required for structured output
output: Output.object({
tag: "calcResult",
schema: z.object({ sum: z.number() }),
}),
});
// output is typed as { sum: number }
console.log("The sum is:", output.sum);
} catch (e) {
if (e instanceof StructuredOutputError) {
console.error("Failed to extract output:", e.rawMatched);
console.error("Cause:", e.cause);
} else {
throw e;
}
}
}
If the agent omits the tag or emits malformed JSON, the catch block captures the StructuredOutputError and exposes the raw text for debugging.
Summary
- Declare the schema using
Output.object({ tag, schema })insrc/Output.tsto establish a type-safe contract with the LLM. - Constrain the run by setting
maxIterations: 1and ensuring the prompt contains the literal opening tag, as enforced insrc/run.ts. - Extract from stdout via
extractStructuredOutput()insrc/extractStructuredOutput.ts, which finds the last tag occurrence, unwraps Markdown fences, parses JSON, and validates against the Standard Schema. - Handle failures by catching
StructuredOutputError, which provides the raw matched text and validation context for debugging.
Frequently Asked Questions
Why must maxIterations be set to 1 for structured output?
The structured output pipeline extracts data after the agent iteration completes, aggregating the entire stdout stream. The library currently only supports this extraction mode for single-turn conversations where maxIterations equals 1. Multi-turn runs would require deterministic extraction between iterations, which the current implementation in src/run.ts does not support (lines 60‑66).
What happens if the agent wraps the JSON in Markdown code fences?
The extraction pipeline automatically handles this case. The unwrapFences() function in src/extractStructuredOutput.ts (lines 50‑58) detects and strips both fenced code blocks (```json ... ```) and plain backtick wrappers before JSON parsing occurs, ensuring the validator receives clean JavaScript objects regardless of formatting.
How do I handle validation errors when extracting structured output?
Wrap your run() call in a try-catch block checking for StructuredOutputError. This error type, defined in src/Output.ts (lines 76‑96), exposes rawMatched (the tag content that failed validation), cause (the parsing or validation error), and run context. This allows you to log the offending snippet, retry with a different prompt, or return a default value.
Can I use validators other than Zod with Output.object()?
Yes. The schema parameter accepts any Standard Schema compatible validator. While the examples use Zod, you can substitute Valibot, ArkType, or other libraries implementing the Standard Schema specification. The validation occurs in src/extractStructuredOutput.ts (lines 57‑80) using the Standard Schema ~validate method, ensuring framework-agnostic type safety.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →