How to Implement Schema-Aligned Parsing for Reliable LLM Outputs
Schema-aligned parsing forces LLMs to emit JSON that validates against a Zod schema, eliminating runtime surprises by catching malformed data before it reaches your business logic.
Unpredictable LLM outputs are a primary source of production bugs in agent systems. The humanlayer/12-factor-agents repository solves this through schema-aligned parsing, a concrete implementation of Factor 4 — "Tools are just structured outputs" that uses the Zod validation library to enforce strict type contracts on model responses.
Why Schema-Aligned Parsing Matters
Without structured validation, LLMs can return ambiguous or malformed data that crashes downstream logic. Schema-aligned parsing provides four critical guarantees:
- Determinism — By forcing the LLM to emit JSON matching a strict schema, you remove ambiguity about data types and structure.
- Safety — Validation catches malformed or unexpected payloads before they affect business logic.
- Observability — Parse errors surface immediately with clear messages, enabling proper logging and monitoring.
- Testability — You can assert that invalid outputs are rejected at the boundary, preventing regressions.
Implementing Schema-Aligned Parsing in 12-Factor Agents
The repository implements a four-step pipeline that turns free-form LLM text into type-safe data structures.
Step 1: Define a Zod Schema for Tool Outputs
Start by declaring the exact shape the LLM must return. In packages/create-12-factor-agent/template/src/a2h.ts, the repository imports z from Zod and defines schemas that serve as the single source of truth for tool arguments.
// src/tools/myTool.ts
import { z } from 'zod';
// The LLM must return an object with `action` (string) and `payload` (object)
export const MyToolSchema = z.object({
action: z.enum(['create', 'delete', 'update']),
payload: z.object({
id: z.string(),
value: z.string().optional(),
}),
});
This schema exports a ZodSchema type that TypeScript can use for compile-time type checking throughout your application.
Step 2: Embed Schema Constraints in Prompts
The LLM needs explicit instructions about the required output format. The repository builds prompts that include the schema description, as seen in packages/walkthroughgen/prompt.md.
// src/promptBuilder.ts
import { MyToolSchema } from './tools/myTool';
export const buildPrompt = (userMessage: string): string => `
You are an assistant. Respond ONLY with JSON that matches this schema:
${MyToolSchema.describe()}
User: ${userMessage}
`;
By embedding schema.describe() in the prompt, you ensure the model understands the exact field types and constraints it must satisfy.
Step 3: Parse and Validate LLM Output
When the LLM returns a response, parse it as JSON and validate against the schema before use. The CLI in packages/walkthroughgen/src/cli.ts demonstrates this pattern, running schema.parse(output) on the raw string.
// src/llmHandler.ts
import { MyToolSchema } from './tools/myTool';
export async function handleUserMessage(msg: string) {
const prompt = buildPrompt(msg);
const raw = await LLM.generate(prompt); // returns a string
try {
const parsed = MyToolSchema.parse(JSON.parse(raw));
// `parsed` is now a typed object you can safely use
return parsed;
} catch (e) {
console.error('Schema validation failed:', e);
throw new Error('Invalid LLM output');
}
}
If validation succeeds, the parsed object is fully typed. If it fails, Zod throws a detailed error describing exactly which constraints were violated.
Step 4: Handle Validation Errors Gracefully
Production agents must react cleanly to malformed outputs. The repository's test suite in packages/walkthroughgen/test/e2e/test-e2e.ts verifies that the system catches bad data early, expecting responses to contain Error: Could not parse YAML content when validation fails.
// test/myTool.test.ts
import { MyToolSchema } from '../src/tools/myTool';
test('rejects malformed JSON', () => {
const bad = '{ "action": "unknown", "payload": { "id": 123 } }';
// id should be string, action should be enum value
expect(() => MyToolSchema.parse(JSON.parse(bad))).toThrow();
});
This guarantees that only schema-compliant data reaches your state machine, as implemented in packages/create-12-factor-agent/template/src/state.ts where the code works with strongly-typed objects like Thread(JSON.parse(data).events).
Production Patterns for Type Safety
Once validation passes, downstream code can assume perfect type safety. The state management logic treats the parsed output as a guaranteed contract, eliminating defensive checks and any types. This pattern aligns with the deep dive on Factor 4 found in content/factor-04-tools-are-structured-outputs.md, which argues that treating tools as structured output contracts makes agents more deterministic and easier to debug.
Summary
Schema-aligned parsing transforms LLM interactions from probabilistic text generation into reliable function calls:
- Define strict contracts using Zod schemas in
packages/create-12-factor-agent/template/src/a2h.ts. - Communicate constraints by embedding schema descriptions in prompts via
packages/walkthroughgen/prompt.md. - Validate every response using
schema.parse()as shown inpackages/walkthroughgen/src/cli.ts. - Catch errors early through automated tests in
packages/walkthroughgen/test/e2e/test-e2e.tsbefore data reaches business logic.
Frequently Asked Questions
What is schema-aligned parsing?
Schema-aligned parsing is the practice of forcing LLMs to return data matching a predefined schema and validating that output before use. According to the humanlayer/12-factor-agents source code, this approach treats tool calls as structured outputs rather than free-form text, eliminating the ambiguity that causes runtime failures in agent systems.
Why does the repository use Zod specifically?
Zod provides runtime validation that matches TypeScript's static type system. As implemented in packages/create-12-factor-agent/template/src/a2h.ts, Zod schemas give you both compile-time type safety and runtime validation with minimal boilerplate, ensuring that data crossing the LLM boundary conforms to expected shapes.
How should you handle LLM outputs that fail schema validation?
The repository's CLI in packages/walkthroughgen/src/cli.ts demonstrates surfacing clear error messages like "Error: Could not parse YAML content" and potentially retrying the request. In production, you should log the raw output for debugging, potentially feed the error back to the LLM for correction, and never propagate untrusted data to downstream functions.
Where does schema-aligned parsing fit within 12-factor agent methodology?
This practice is the concrete implementation of Factor 4: "Tools are just structured outputs" documented in content/factor-04-tools-are-structured-outputs.md. By treating every tool call as a contract with a strict input/output schema, agents become deterministic, observable, and testable—core requirements for production-grade AI systems.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →