How to Use the wren-pydantic SDK for Type-Safe WrenAI Integration

The wren-pydantic SDK converts Wren CLI-prepared projects into Pydantic AI toolkits, enabling fully typed SQL generation and validation through WrenToolkit.from_project() and toolkit.toolset().

The wren-pydantic SDK bridges WrenAI's semantic layer with Pydantic AI's agent framework, allowing developers to build type-safe LLM applications that query databases through natural language. As implemented in the Canner/WrenAI repository, this integration transforms compiled MDL (Model Definition Language) artifacts into executable Python tools with strict Pydantic validation.

Architecture Overview

The SDK follows a five-stage pipeline that maintains type safety from project configuration to agent execution.

Project Bootstrap

The Wren CLI initializes a project directory containing wren_project.yml, the compiled target/mdl.json, connection profiles, and optional memory indices. These artifacts serve as the single source of truth for schema definitions and connection parameters.

Toolkit Construction

In sdk/wren-pydantic/src/wren_pydantic/_toolkit.py, the WrenToolkit.from_project() class method consumes these CLI artifacts to instantiate a WrenToolkit object. This object exposes high-level methods including query(), dry_plan(), and memory.fetch(), while internally managing SQL execution engines and schema context retrieval.

Toolset Generation

The toolkit.toolset() method, defined in sdk/wren-pydantic/src/wren_pydantic/_tools.py, generates a Pydantic AI FunctionToolset. This exposes LLM-facing tools such as wren_query, wren_fetch_context, and wren_recall_queries, each configured with retries=2 and automatic error mapping to ModelRetry or WrenError exceptions.

Instruction Synthesis

toolkit.instructions() dynamically generates system prompts that adapt to the enabled tools, instructing the model on the exact workflow: recall historical queries → fetch relevant context → execute SQL → store results for future retrieval.

Type-Safe Execution

All tool returns use Pydantic models defined in sdk/wren-pydantic/src/wren_pydantic/_models.py (WrenQueryResult, FetchContextResult, etc.), enabling direct validation of LLM outputs against typed schemas.

Installation and Setup

Ensure you have a Wren CLI-prepared project directory before initializing the SDK. The SDK is deliberately synchronous; Pydantic AI handles the async bridging internally.

pip install wren-pydantic pydantic-ai

Building Type-Safe Agents

Basic Agent Setup

Attach to an existing Wren project and create an agent with the full toolset:

from wren_pydantic import WrenToolkit
from pydantic_ai import Agent

# Load project artifacts from CLI output

toolkit = WrenToolkit.from_project("./analytics_db")

# Initialize agent with generated instructions and tools

agent = Agent(
    "openai:gpt-4o",
    instructions=toolkit.instructions(),
    toolsets=[toolkit.toolset()],
)

# Execute natural language query

result = agent.run_sync("Top 5 customers by revenue last quarter?")
print(result.output)

This configuration includes memory write tools by default, allowing the agent to store query history via wren_store_query.

Structured Output Validation

Enforce type safety on final outputs using Pydantic models:

from pydantic import BaseModel
from wren_pydantic import WrenToolkit
from pydantic_ai import Agent

class TopCustomers(BaseModel):
    period: str
    customers: list[str]

toolkit = WrenToolkit.from_project("./analytics_db")

agent = Agent(
    "openai:gpt-4o",
    instructions=toolkit.instructions(),
    toolsets=[toolkit.toolset()],
    output_type=TopCustomers,  # Validates against schema

)

res = agent.run_sync("Top 5 customers last quarter?")
print(res.output.customers)  # Typed as list[str]

The agent validates the LLM's final response against TopCustomers before returning, ensuring programmatic reliability.

Read-Only Memory Configuration

For shared or curated projects where agents should not modify the memory index, disable write operations:

toolkit = WrenToolkit.from_project("./public_project")
toolset = toolkit.toolset(include_memory_write=False)  # Excludes wren_store_query

agent = Agent(
    "openai:gpt-4o",
    instructions=toolkit.instructions(toolset=toolset),
    toolsets=[toolset],
)

This pattern, implemented in sdk/wren-pydantic/src/wren_pydantic/_tools_memory.py, removes the wren_store_query tool while preserving read capabilities like wren_fetch_context and wren_recall_queries.

Integrating with Custom Dependencies

Combine Wren tools with application-specific dependencies using deps_type:

from dataclasses import dataclass
from pydantic_ai import Agent, RunContext
from wren_pydantic import WrenToolkit

@dataclass
class MyDeps:
    api_client: ApiClient

toolkit = WrenToolkit.from_project("./analytics_db")

agent = Agent(
    "openai:gpt-4o",
    deps_type=MyDeps,
    toolsets=[toolkit.toolset(takes_ctx=True)],  # Required for context injection

)

@agent.tool
def lookup_external(ctx: RunContext[MyDeps], id: str) -> str:
    return ctx.deps.api_client.fetch(id)

Setting takes_ctx=True in toolkit.toolset() enables the Wren tools to access the RunContext, allowing seamless integration with external APIs alongside database queries.

Multi-Project Orchestration

Instantiate multiple toolkits within a single program to query disparate data domains:

loans = WrenToolkit.from_project("./loans_proj")
events = WrenToolkit.from_project("./events_proj")

loans_agent = Agent(
    "openai:gpt-4o",
    instructions=loans.instructions(),
    toolsets=[loans.toolset()],
)

events_agent = Agent(
    "openai:gpt-4o",
    instructions=events.instructions(),
    toolsets=[events.toolset()],
)

# Direct toolkit access for manual orchestration

loans_results = loans.query("SELECT * FROM loan")
events_results = events.query("SELECT * FROM event")

Each WrenToolkit maintains independent connection profiles and memory indices, preventing cross-contamination between projects.

Error Handling and Validation

The SDK provides specific exceptions in sdk/wren-pydantic/src/wren_pydantic/exceptions.py:

  • WrenToolkitInitError: Raised when project artifacts are missing or corrupted
  • MemoryNotEnabledError: Occurs when memory-dependent tools are invoked on toolkit instances without LanceDB configuration

Error conversion utilities (to_model_retry, should_propagate) automatically map transient failures to Pydantic AI's retry mechanisms, preventing agent stalls while surfacing critical configuration errors immediately.

Summary

  • wren-pydantic transforms Wren CLI projects into Pydantic AI toolkits via WrenToolkit.from_project().
  • Type safety is enforced through Pydantic models in _models.py that validate both tool inputs and agent outputs.
  • The toolset() method generates configurable FunctionToolset instances with optional memory write capabilities.
  • Integration supports custom dependencies through takes_ctx=True and multi-project isolation through separate toolkit instances.
  • Error handling distinguishes between retryable transient failures and critical initialization errors.

Frequently Asked Questions

What prerequisites are required for the wren-pydantic SDK?

You must first create a Wren project using the Wren CLI, which generates the wren_project.yml, target/mdl.json, and connection profiles required by WrenToolkit.from_project(). The SDK does not compile MDL definitions itself; it consumes pre-built artifacts.

How does the SDK handle database connections?

The WrenToolkit class reads connection profiles from the project directory and manages SQL execution engines internally. According to sdk/wren-pydantic/src/wren_pydantic/_toolkit.py, it exposes synchronous methods like query() and dry_plan() that Pydantic AI wraps for async execution, ensuring no hidden concurrency issues.

Can I disable the memory features entirely?

Yes. If the project directory lacks a LanceDB index or you explicitly set include_memory_write=False in toolkit.toolset(), the memory tools (wren_fetch_context, wren_recall_queries, wren_store_query) are excluded from the generated toolset. The agent will still execute SQL through wren_query but without semantic memory of past interactions.

What is the difference between instructions() and toolset()?

toolkit.toolset() returns the executable FunctionToolset containing the actual Python functions the LLM can call, while toolkit.instructions() generates a system prompt string describing the workflow and available tools. You must provide both to the Agent constructor for the integration to function correctly, as shown in sdk/wren-pydantic/src/wren_pydantic/_toolkit.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →