How LangGraph Transformations Process Source Content in Open-Notebook
Open-Notebook implements content transformations as a compiled LangGraph workflow that extracts source text, applies customizable prompt templates, invokes provisioned LLMs, and persists cleaned results as source insights.
Open-Notebook leverages LangGraph to orchestrate content-wise transformations through a compact, deterministic workflow defined in open_notebook/graphs/transformation.py. This architecture processes raw source material—whether passed directly as text or retrieved from persisted records—through a standardized pipeline of prompt construction, model invocation, and result persistence. Understanding how these LangGraph transformations handle source content reveals the system's model-agnostic approach to AI-powered document analysis.
Defining the Transformation State
The workflow relies on TransformationState, a TypedDict defined in open_notebook/graphs/transformation.py that tracks the complete execution context. This state holds four critical components:
input_text– The raw text submitted for transformation (optional if aSourceis provided).source– ASourceobject containing the full extracted text and metadata.transformation– TheTransformationdefinition including the prompt template and title.output– The final cleaned result produced by the LLM.
This state structure ensures that the graph can operate in two modes: processing ad-hoc text supplied directly, or processing persisted content from the notebook's source repository.
The Transformation Execution Flow
The LangGraph workflow consists of a single node (run_transformation) that executes a six-step pipeline from text ingestion to insight persistence.
Text Retrieval and State Preparation
The run_transformation node first extracts the Source object and input content from the state. If input_text is absent, the node automatically falls back to source.full_text—the complete extracted text of the persisted source record. This fallback mechanism guarantees that the workflow functions whether the caller supplies temporary text or references a stored document.
Prompt Construction and Model Invocation
Once the raw content is secured, the system constructs the final prompt through three layers:
- Default Instructions: The system retrieves global default prompts from
DefaultPrompts(defined inopen_notebook/domain/transformation.py) and optionally prefixes them to the transformation-specific template. - Template Rendering: The
transformation.prompttemplate is rendered with the current state usingai_prompter.Prompter, producing the final instruction set. - Message Construction: The rendered prompt is wrapped in a
SystemMessage, while the source content is sent as aHumanMessage.
The provisioned model is instantiated via provision_langchain_model (located in open_notebook/ai/provision.py), which builds a LangChain chain targeting the model ID specified in the LangGraph config (configurable["model_id"]). The chain is invoked asynchronously using chain.ainvoke(payload).
Response Processing and Insight Persistence
The LLM's raw response.content undergoes two sanitization steps before storage:
extract_text_contentstrips the response to plain text, removing any metadata wrappers.clean_thinking_contentremoves chain-of-thought markers or "thinking" meta-information that models like Claude or DeepSeek may emit.
If a Source instance is present in the state, the cleaned result is persisted as an insight using await source.add_insight(transformation.title, cleaned_content), linking the transformation output directly to the source document.
Graph Structure and API Integration
The transformation graph uses minimal wiring: a single node named "agent" connected from START to END. The compiled graph (graph = agent_state.compile()) is invoked from the API endpoint POST /transformations/execute in api/routers/transformations.py.
The endpoint accepts the input_text, selected Transformation record, and model_id, then returns the transformed output after the graph completes execution.
Practical Implementation Examples
Direct Graph Invocation (Python)
Invoke the transformation pipeline programmatically by compiling the graph with your source and model configuration:
from open_notebook.graphs.transformation import graph as transformation_graph
from open_notebook.domain.transformation import Transformation
from open_notebook.domain.notebook import Source
# Assume a Source object already exists
source = await Source.get(source_id)
# Load a transformation definition (e.g., "Summarize")
transformation = await Transformation.get(transformation_id)
# Run the graph, providing the source and choosing a model
result = await transformation_graph.ainvoke(
{
"source": source, # The Source object (optional if input_text supplied)
"transformation": transformation,
},
config={"configurable": {"model_id": "openai:gpt-4o"}}, # model selection
)
print("Transformed output:", result["output"])
Public API Execution (cURL)
Trigger transformations via the REST API when operating the Open-Notebook server:
curl -X POST https://localhost:5055/transformations/execute \
-H "Content-Type: application/json" \
-d '{
"transformation_id": "summarize",
"model_id": "anthropic:claude-3.5-sonnet",
"input_text": "Long article text ..."
}'
Response:
{
"output": "A concise summary of the article …",
"translation_id": "summarize",
"model_id": "anthropic:claude-3.5-sonnet"
}
Customizing Default Instructions (cURL)
Modify the system-wide default prompts that prefix all transformation templates:
curl -X PUT https://localhost:5055/transformations/default-prompt \
-H "Content-Type: application/json" \
-d '{ "transformation_instructions": "Always keep the original tone." }'
Key Implementation Files
| Component | File Path | Role |
|---|---|---|
| Transformation Graph | open_notebook/graphs/transformation.py |
Defines TransformationState, run_transformation node, and graph wiring. |
| Transformation Model | open_notebook/domain/transformation.py |
Pydantic schemas for Transformation records and DefaultPrompts configuration. |
| Source Model | open_notebook/domain/notebook.py |
Provides full_text property and add_insight method for persistence. |
| LLM Provisioning | open_notebook/ai/provision.py |
Implements provision_langchain_model for model-agnostic chain creation. |
| API Router | api/routers/transformations.py |
Exposes POST /transformations/execute endpoint. |
| Text Utilities | open_notebook/utils/text_utils.py (extract), open_notebook/utils/__init__.py (cleaning) |
Utilities for normalizing LLM output before persistence. |
Summary
- LangGraph transformations in Open-Notebook follow a single-node workflow defined in
open_notebook/graphs/transformation.py. - The system automatically falls back to
source.full_textwheninput_textis not provided, enabling both ad-hoc and persisted content processing. - Prompts combine
DefaultPrompts(global instructions) with transformation-specific templates before being sent to provisioned models viaprovision_langchain_model. - LLM responses are sanitized using
extract_text_contentandclean_thinking_contentto remove meta-information and formatting artifacts. - Results are persisted as source insights using
source.add_insight, creating a permanent link between the transformation output and the original document.
Frequently Asked Questions
What happens if I don't provide input_text to a transformation?
The run_transformation node checks for input_text in the state and automatically falls back to source.full_text when the direct input is missing. This ensures the workflow processes the full extracted text of any persisted Source object, making the transformation flexible for both temporary text and stored documents.
How does the system handle different LLM providers?
The provision_langchain_model function in open_notebook/ai/provision.py dynamically constructs a LangChain chain based on the model_id passed in the LangGraph configuration (configurable["model_id"]). This allows the same graph to execute using OpenAI, Anthropic, or other compatible providers without code changes, requiring only the model identifier in the format "provider:model-name".
Where are transformation results stored?
When a Source object is present in the graph state, the cleaned transformation output is persisted via await source.add_insight(transformation.title, cleaned_content). This attaches the result as an insight to the source document, making it available for future reference within the notebook interface. If no source is provided, the result is returned directly in the graph output without persistence.
Can I customize the system prompts for all transformations?
Yes. The DefaultPrompts class in open_notebook/domain/transformation.py stores global instructions that are optionally prefixed to every transformation-specific prompt. You can update these defaults via the PUT /transformations/default-prompt API endpoint, allowing you to inject persistent instructions—such as tone guidelines or formatting requirements—across all transformation executions.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →