How Source Transformations Are Processed Through the LangGraph Workflow in Open-Notebook
Open-Notebook implements source transformations as a compiled LangGraph workflow that extracts content, applies templated prompts, provisions model-specific chains, and persists cleaned results as source insights.
The open-notebook repository orchestrates content transformation through a structured state machine built on LangGraph. This architecture ensures that source transformations processed through the LangGraph workflow follow a deterministic pipeline from raw text extraction to insight generation. The implementation leverages asynchronous nodes and configurable model provisioning to handle diverse transformation tasks across different AI providers.
LangGraph Workflow Architecture
State Definition with TransformationState
The workflow state is defined in open_notebook/graphs/transformation.py via the TransformationState TypedDict. This state container holds four critical fields:
input_text: Raw text supplied directly to the graph invocationsource: TheSourceobject containing persisted content and metadatatransformation: TheTransformationdefinition including prompt templatesoutput: The final transformed string produced by the LLM
The run_transformation Node Implementation
The core processing logic resides in the run_transformation node. This function executes the following sequence:
-
Content Resolution: Extracts the
Sourceobject and checks forinput_text. If the input text is missing, it falls back tosource.full_text(the full extracted text of the persisted source), ensuring the workflow functions with both ad-hoc text and database records. -
Prompt Construction: Retrieves the transformation's prompt template (
transformation.prompt) from the state. It optionally prefixes this with system-wide default instructions stored inDefaultPrompts(defined inopen_notebook/domain/transformation.py). -
Message Preparation: Renders the final prompt using
ai_prompter.Prompterwith the current state as context. The rendered prompt is wrapped in aSystemMessage, while the raw source content is sent as aHumanMessage.
Model Provisioning and Asynchronous Execution
The graph provisions language models through provision_langchain_model in open_notebook/ai/provision.py. This utility:
- Reads the
model_idfrom the LangGraph runtime configuration (configurable["model_id"]) - Constructs a LangChain chain targeting the specific model provider (e.g., OpenAI, Anthropic)
- Invokes the chain asynchronously via
chain.ainvoke(payload)
This design enables model-agnostic transformations where the same compiled graph executes against different providers based on runtime configuration, eliminating hardcoded model dependencies.
Response Cleaning and Insight Persistence
After LLM execution, the workflow handles post-processing before persistence:
-
Content Extraction: The raw
response.contentis normalized to plain text usingextract_text_contentfromopen_notebook/utils/text_utils.py. -
Thinking Content Removal:
clean_thinking_content(defined inopen_notebook/utils/__init__.py) strips meta-information such as chain-of-thought markers that some reasoning models emit. -
Persistence: If a
Sourceinstance is present in the state, the cleaned result is saved viaawait source.add_insight(transformation.title, cleaned_content), storing the transformation as a structured insight attached to the source object.
Graph Structure and API Integration
The graph wiring in open_notebook/graphs/transformation.py defines a minimal single-node flow: START → "agent" (the run_transformation node) → END. The compiled graph (graph = agent_state.compile()) is exposed through the REST API endpoint POST /transformations/execute in api/routers/transformations.py.
The endpoint accepts:
transformation_id: The transformation definition to applymodel_id: Target model identifier (e.g.,"openai:gpt-4o")input_text(optional): Ad-hoc text or null to use the source's stored content
Practical Implementation Examples
Direct Graph Invocation (Python)
from open_notebook.graphs.transformation import graph as transformation_graph
from open_notebook.domain.transformation import Transformation
from open_notebook.domain.notebook import Source
# Load existing source and transformation definitions
source = await Source.get(source_id)
transformation = await Transformation.get(transformation_id)
# Execute with specific model configuration
result = await transformation_graph.ainvoke(
{
"source": source,
"transformation": transformation,
},
config={"configurable": {"model_id": "openai:gpt-4o"}},
)
print(result["output"])
REST API Call
curl -X POST http://localhost:5055/transformations/execute \
-H "Content-Type: application/json" \
-d '{
"transformation_id": "summarize",
"model_id": "anthropic:claude-3.5-sonnet",
"input_text": "Long article text ..."
}'
Response:
{
"output": "A concise summary of the article …",
"transformation_id": "summarize",
"model_id": "anthropic:claude-3.5-sonnet"
}
Configuring Default Prompts
Global transformation instructions can be updated via:
curl -X PUT http://localhost:5055/transformations/default-prompt \
-H "Content-Type: application/json" \
-d '{ "transformation_instructions": "Always keep the original tone." }'
Summary
- Open-Notebook uses a single-node LangGraph workflow defined in
open_notebook/graphs/transformation.pyto handle all source transformations - The
TransformationStateinterface manages the flow between raw input, transformation definitions, and model output - Content extraction follows a fallback pattern: explicit
input_texttakes precedence oversource.full_text - Model provisioning is dynamic via
provision_langchain_model, supporting multiple providers through runtime configuration - Post-processing removes model meta-content and persists results as source insights using
add_insight
Frequently Asked Questions
What triggers the LangGraph transformation workflow?
The workflow is triggered by calling the compiled graph's ainvoke method, either directly in Python or through the POST /transformations/execute API endpoint. The caller must provide a Transformation definition and optionally a Source object or raw input_text.
How does the workflow handle missing input text?
If the input_text field is absent from the state, the run_transformation node automatically falls back to source.full_text. This ensures the graph works with both ephemeral text inputs and persisted source records from the database.
Where are transformation prompts configured?
Prompt templates are stored in the Transformation model (defined in open_notebook/domain/transformation.py). The system also supports global default instructions via DefaultPrompts, which are automatically prefixed to every transformation prompt before sending to the LLM.
How does Open-Notebook support multiple AI models in the same transformation graph?
The graph accepts a model_id in its configurable runtime parameters. The provision_langchain_model function in open_notebook/ai/provision.py dynamically constructs the appropriate LangChain chain for the specified model (e.g., "openai:gpt-4o" or "anthropic:claude-3.5-sonnet"), enabling model-agnostic transformation execution.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →