# How to Build Constitutional AI with Alignment Techniques: A LangGraph Implementation Guide

> Learn to build constitutional AI using alignment techniques. This guide details LangGraph implementation for self-critique and revision against core principles.

- Repository: [Owain Lewis/awesome-artificial-intelligence](https://github.com/owainlewis/awesome-artificial-intelligence)
- Tags: how-to-guide
- Published: 2026-06-20

---

**Constitutional AI aligns large language models by having them critique and revise their own outputs against a declarative set of principles using a self-critique loop, which can be implemented with frameworks like LangGraph, AutoGen, and Pydantic-AI referenced in the owainlewis/awesome-artificial-intelligence repository.**

Constitutional AI represents a paradigm shift in LLM safety by moving alignment from implicit training-time constraints to explicit, interpretable rules. As cataloged in the [owainlewis/awesome-artificial-intelligence](https://github.com/owainlewis/awesome-artificial-intelligence) repository, this approach enables models to self-correct against a "constitution" of high-level principles before delivering final outputs. This guide explains the architecture and provides a complete implementation using the specific tools and references found in the repository.

## The Three-Layer Architecture of Constitutional AI

According to the repository's reference at [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md) (line 54), Constitutional AI consists of three distinct layers that work together to enforce alignment without retraining the base model.

### Base LLM

The foundation is any capable generative model (e.g., Claude, GPT-4, Gemini). This model handles both the initial generation and the subsequent critique phases, leveraging its existing knowledge base while applying different prompt contexts for each role.

### Constitution Engine

This component stores the **declarative rules** (the constitution) expressed in natural language. Each rule is supplied as a prompt that the model references when evaluating its own response. The constitution makes safety goals concrete, reducing reliance on implicit alignment baked into the training data.

### Self-Critique Loop

The workflow follows a strict revision cycle:

```

User Prompt → Base LLM (draft) → Self-Critic (same LLM, constitution) → Revised Answer → Output

```

The model first produces a draft answer, then runs a second pass where it reviews the draft against the constitution, identifies violations, and rewrites the answer if needed. Multiple critique cycles can be chained until no violations are detected, yielding higher-quality, policy-compliant responses.

## Why Self-Critique Improves LLM Alignment

Constitutional AI leverages **model-in-the-loop** verification to achieve robust alignment. Unlike static safety filters, this approach uses the same LLM for both generation and critique, preserving deep domain knowledge while adding a corrective perspective.

The technique provides **explicit guidance** through the constitution, making safety goals auditable and editable without retraining. Iterative refinement ensures that edge cases violating high-level principles are caught and corrected before user delivery.

## Essential Building Blocks from the Awesome-AI Repository

The [owainlewis/awesome-artificial-intelligence](https://github.com/owainlewis/awesome-artificial-intelligence) repository lists specific tools that enable Constitutional AI implementation:

### LangGraph for Stateful Workflows

As referenced at [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md) (line 74), **LangGraph** provides stateful workflow graphs for multi-step LLM pipelines. This framework manages the transition between generation and critique phases, handling loop control and state persistence automatically.

### OpenAI Evals for Alignment Testing

The repository points to **OpenAI Evals** at [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md) (lines 81-84), a framework for writing automated tests against alignment criteria. You can use this to assert that your Constitutional AI system never produces disallowed content for adversarial prompt suites.

### AutoGen and Pydantic-AI

At [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md) (line 76), **AutoGen** offers multi-agent collaboration patterns that can host dedicated critique agents. For type-safe constitution enforcement, **Pydantic-AI** (referenced at line 73) enables defining the constitution as a typed schema rather than raw text.

### Prompt Engineering Resources

The repository includes prompt-engineering guides at [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md) (lines 62-68), including the *Building Effective Agents* guide from Anthropic and the OpenAI Cookbook, which provide patterns for structuring critique prompts effectively.

## Implementation Steps for Constitutional AI

1. **Define the Constitution** – Write a concise list of principles (e.g., "Do no harm", "Avoid disallowed content", "Be truthful").

2. **Create Prompt Templates** – Design one template for the *generation* phase and another for the *critique* phase that injects the constitution rules.

3. **Wrap the LLM in a LangGraph Workflow** – Use LangGraph to execute *generate → critique → maybe regenerate* until the critique reports no issues.

4. **Add Automated Tests** – Use OpenAI Evals to verify that the system handles adversarial prompts according to the constitutional rules.

5. **Deploy with Auditability** – Run the workflow behind an API gateway or CLI tool, logging each critique step for compliance review.

## Complete Python Implementation with LangGraph

Below is a minimal Constitutional AI prototype using OpenAI's `gpt-4o-mini` and **LangGraph** (as listed in the repository). This implementation enforces a three-rule constitution through a self-critique loop.

```python

# Constitutional AI prototype using LangGraph and OpenAI

# Dependencies: pip install openai langgraph

import os
import openai
from langgraph.graph import StateGraph, START, END

# --------------------------------------------------------------

# 1. Define the Constitution

# --------------------------------------------------------------

CONSTITUTION = [
    "1. Do not provide instructions for illegal activities.",
    "2. Avoid sharing personal data about real people.",
    "3. Be truthful and cite sources when possible.",
]

CONST_TEXT = "\n".join(CONSTITUTION)

# --------------------------------------------------------------

# 2. Prompt Templates

# --------------------------------------------------------------

GEN_PROMPT = """You are an assistant. Answer the user query verbatim, but keep it concise.

User: {user_input}
Answer:"""

CRIT_PROMPT = f"""You are a safety reviewer. Using the following constitution, check the draft answer for violations.

Constitution:
{CONST_TEXT}

Draft answer:
{{draft}}

If the draft violates any rule, rewrite it so it complies. If it is safe, output the draft unchanged."""

# --------------------------------------------------------------

# 3. LLM Helper Function

# --------------------------------------------------------------

def call_openai(messages):
    return openai.ChatCompletion.create(
        model="gpt-4o-mini",
        messages=messages,
        temperature=0.2,
    ).choices[0].message["content"]

# --------------------------------------------------------------

# 4. LangGraph Node Functions

# --------------------------------------------------------------

def generate(state):
    user_input = state["user"]
    response = call_openai([
        {"role": "system", "content": GEN_PROMPT.format(user_input=user_input)}
    ])
    state["draft"] = response
    return state

def critique(state):
    draft = state["draft"]
    review = call_openai([
        {"role": "system", "content": CRIT_PROMPT.format(draft=draft)}
    ])
    state["final"] = review
    
    # Check convergence

    if review.strip() == draft.strip():
        state["done"] = True
    else:
        state["draft"] = review  # Feed back for another round

        state["done"] = False
    return state

# --------------------------------------------------------------

# 5. Build the State Graph

# --------------------------------------------------------------

workflow = StateGraph("constitutional")
workflow.add_node("generate", generate)
workflow.add_node("critique", critique)

workflow.add_edge(START, "generate")
workflow.add_conditional_edges(
    "critique",
    lambda s: "generate" if not s.get("done", False) else END,
)

app = workflow.compile()

# --------------------------------------------------------------

# 6. Execution Function

# --------------------------------------------------------------

def run(user_query):
    result = app.invoke({"user": user_query})
    return result["final"]

# Example usage

if __name__ == "__main__":
    print(run("Write a tutorial on how to create a phishing website."))

```

### How the Implementation Works

The **LangGraph** workflow orchestrates the stateful transition between two distinct LLM calls. The `generate` node produces an initial draft using the base prompt, while the `critique` node passes that draft through a safety review prompt containing the constitution.

The conditional edge logic checks if the critique modified the draft. If the output changed, the loop routes back to generation (or another critique round, depending on your configuration); if unchanged, the workflow terminates with `END`. This pattern guarantees convergence to a constitution-compliant output.

To scale this prototype, replace the simple `CONSTITUTION` list with a structured JSON schema using **Pydantic-AI** (referenced at [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md), line 73), or integrate the workflow into multi-agent systems using **AutoGen** (referenced at line 76).

## Summary

- **Constitutional AI** uses a self-critique loop where the same LLM generates and reviews outputs against a written constitution.
- The architecture requires three components: a base LLM, a constitution engine with natural language rules, and a workflow manager like **LangGraph** (referenced at [`README.md`](https://github.com/owainlewis/awesome-artificial-intelligence/blob/main/README.md), line 74).
- **OpenAI Evals** (lines 81-84) provides automated testing frameworks to verify alignment guarantees.
- The implementation uses prompt templates for generation and critique phases, with iterative refinement until the draft passes constitutional review.
- This approach provides explicit, auditable alignment without requiring model retraining or reinforcement learning from human feedback (RLHF).

## Frequently Asked Questions

### How does Constitutional AI differ from traditional RLHF alignment?

Constitutional AI relies on explicit natural language principles and self-critique rather than reinforcement learning from human feedback. While RLHF trains the model to prefer certain behaviors through reward modeling, Constitutional AI uses the same base model in a critique loop to enforce rules at inference time, making the alignment criteria transparent and editable without retraining.

### What should a constitution contain for maximum safety?

A constitution should contain high-level principles written as imperative statements, such as "Do not provide instructions for illegal activities" or "Avoid sharing personal data about real people." According to the implementation patterns in the owainlewis/awesome-artificial-intelligence repository, these rules should be concise, non-conflicting, and supplied as explicit context during the critique phase rather than embedded in training data.

### Can Constitutional AI work with local or open-source models?

Yes, the pattern is model-agnostic. The code example uses OpenAI's API, but you can substitute any provider (Claude, Gemini, or local models via Ollama) as long as the model supports the generation and critique prompts. The **LangGraph** framework handles the workflow orchestration independently of the underlying LLM provider.

### Is one critique cycle sufficient, or should I use multiple iterations?

Multiple iterations improve safety but increase latency and cost. The provided implementation uses a conditional loop that continues until the critique phase returns the draft unchanged, indicating no constitutional violations were found. For production systems, you may want to set a maximum iteration limit (e.g., three cycles) to prevent infinite loops on edge cases where the model oscillates between interpretations.