How to Implement Text Encoding with CLIP and Dynamic Prompts in ComfyUI
To implement text encoding with CLIP and dynamic prompts in ComfyUI, load a CLIP model using the CLIPLoader node, then pass the model and a text string to the CLIPTextEncode node, ensuring the text input is marked with dynamicPrompts=True to enable runtime placeholder expansion.
ComfyUI provides a modular architecture for Stable Diffusion workflows where text encoding with CLIP and dynamic prompts enables runtime prompt modification without changing the node graph structure. This guide examines the source code implementation in the Comfy-Org/ComfyUI repository to show how the CLIPTextEncode node processes text with CLIP models and how the execution engine resolves placeholders before tokenization.
CLIP Text Encoding Architecture
The implementation involves three core components that work together to transform text into conditioning tensors:
-
CLIPLoader(nodes.pylines 975-1002): Loads a CLIP or CLIP-Vision checkpoint and makes the model available to downstream nodes via theCLIPoutput type. -
CLIPTextEncode(nodes.pylines 59-81): Accepts aCLIPmodel and atextstring, tokenizes the input usingclip.tokenize(text), and returns a conditioning tensor viaclip.encode_from_tokens_scheduledthat samplers consume. -
DynamicPrompt Engine (
comfy_execution/graph.pylines 22-38): Wraps the prompt dictionary and expands placeholders like${artist}before the node executes, storing the original prompt and an ephemeral copy for substitutions.
The execution flow follows this sequence: First, CLIPLoader produces a CLIP model instance. Next, CLIPTextEncode receives this model along with a text input. If the text input declares "dynamicPrompts": True in its definition (as seen in comfy_api/latest/_io.py lines 25-34), the engine creates a DynamicPrompt wrapper when the workflow submits (as implemented in execution.py lines 704-712). The wrapper expands any placeholders using variables from the extra_data dictionary, then the node tokenizes the final string.
How Dynamic Prompts Work
The dynamic prompt system operates through three distinct layers that handle declaration, serialization, and runtime expansion.
Widget Declaration
In the node definition, the text input specifies dynamicPrompts=True within its input options. This flag, defined in comfy_api/latest/_io.py, instructs the frontend that the field may contain Jinja-style expressions or placeholder variables.
Runtime Expansion
The DynamicPrompt class in comfy_execution/graph.py maintains two internal dictionaries:
self.original_prompt # Raw JSON received from the client
self.ephemeral_prompt # Mutable copy where dynamic text is substituted
When execution.py processes a workflow (lines 704-712), it instantiates this wrapper. As the topological sorter executes each node, the engine evaluates placeholders against the extra_data dictionary passed in the API request or built-in helpers like ${prompt}.
Evaluation Logic
The expansion logic resides in comfy/prompt_parser.py (invoked by the DynamicPrompt wrapper). This parser resolves syntax such as ${variable}, ${choose|option1|option2}, and ${if:condition|true text|false text} before the text ever reaches the CLIP tokenizer.
Dynamic Prompt Usage Patterns
Depending on your workflow requirements, you can implement several patterns for text encoding:
-
Static prompts: Provide a plain string such as
"a portrait of a cyberpunk city"with no special syntax required. -
Variable substitution: Pass variables in the
extra_datapayload and reference them as"A painting by ${artist}". The engine substitutes values before tokenization. -
Random selection: Use the built-in
${choose|option1|option2|option3}syntax to randomly select alternatives during each execution. -
Conditional blocks: Implement branching logic with
${if:condition|true text|false text}to modify prompts based on runtime conditions.
All evaluations occur once per execution, and the resulting final string feeds into CLIPTextEncode.
Implementation Examples
Static Prompt Workflow
The minimal JSON workflow demonstrates a basic text encoding without dynamic elements:
{
"1": {
"class_type": "CLIPLoader",
"inputs": {
"ckpt_name": "clip-vit-large-patch14.safetensors"
}
},
"2": {
"class_type": "CLIPTextEncode",
"inputs": {
"clip": ["1", 0],
"text": "A futuristic city at sunset"
}
}
}
The text field contains a literal string that the node tokenizes directly without expansion.
Dynamic Prompt with Placeholders
To substitute variables at runtime, include extra_data in your payload:
{
"extra_data": {
"artist": "Studio Ghibli"
},
"1": {
"class_type": "CLIPLoader",
"inputs": { "ckpt_name": "clip-vit-large-patch14.safetensors" }
},
"2": {
"class_type": "CLIPTextEncode",
"inputs": {
"clip": ["1", 0],
"text": "A beautiful landscape in the style of ${artist}"
}
}
}
When executed, the engine replaces ${artist} with Studio Ghibli before the CLIPTextEncode node processes the text.
Python API Submission
You can programmatically submit workflows with dynamic prompts using the ComfyUI API:
import json, urllib.request
# Build the prompt JSON with variables
prompt = {
"extra_data": {"artist": "Alex Ross"},
"1": {"class_type": "CLIPLoader",
"inputs": {"ckpt_name": "clip-vit-large-patch14.safetensors"}},
"2": {"class_type": "CLIPTextEncode",
"inputs": {"clip": ["1", 0],
"text": "A heroic portrait in the style of ${artist}"}}
}
# POST to the ComfyUI API
data = json.dumps({"prompt": prompt}).encode()
req = urllib.request.Request("http://127.0.0.1:8188/prompt", data=data)
urllib.request.urlopen(req) # Execution is queued
The request includes extra_data, so the server resolves ${artist} to Alex Ross before the CLIP node receives the final string.
Custom Node with Dynamic Support
To create a custom node that supports dynamic prompts, declare the input with the dynamicPrompts flag:
class MyPromptNode:
@classmethod
def INPUT_TYPES(s):
return {
"required": {
"text": (
"STRING",
{"multiline": True, "dynamicPrompts": True,
"tooltip": "Enter a prompt; can contain ${variables}."}
)
}
}
RETURN_TYPES = ("STRING",)
FUNCTION = "process"
def process(self, text):
# text is already the expanded string
return (text,)
Because the input definition mirrors CLIPTextEncode, the engine automatically expands placeholders before calling the process method.
Key Source Files
Understanding these specific source files helps when debugging or extending the functionality:
-
nodes.py(lines 59-81): Contains theCLIPTextEncodeclass that tokenizes and encodes text with CLIP models. -
nodes.py(lines 975-1002): ImplementsCLIPLoader, which supplies the CLIP model used by text encoding nodes. -
comfy_api/latest/_io.py(lines 25-34): Defines theStringwidget type and thedynamic_promptsflag that serializes into the JSON payload. -
comfy_execution/graph.py(lines 22-38): Houses theDynamicPromptwrapper class that manages prompt expansion at runtime. -
execution.py(lines 704-712): Handles the creation ofDynamicPromptobjects for incoming workflow requests. -
script_examples/basic_api_example.py(lines 53-71): Demonstrates client-side workflow construction and API submission patterns.
Summary
Implementing text encoding with CLIP and dynamic prompts in ComfyUI requires understanding the interaction between model loading, node execution, and runtime text expansion:
- Load CLIP models using
CLIPLoaderfromnodes.pybefore any text encoding operations. - Use
CLIPTextEncodeto tokenize text and generate conditioning tensors for samplers. - Enable dynamic prompts by setting
dynamicPrompts=Truein the input definition, allowing the engine to resolve${variables}before tokenization. - Pass substitution variables via the
extra_datafield in API requests to modify prompts at runtime without changing the workflow graph. - Reference
comfy_execution/graph.pyandexecution.pyto understand how theDynamicPromptwrapper manages prompt expansion.
Frequently Asked Questions
How does the CLIPTextEncode node tokenize input text?
The CLIPTextEncode node in nodes.py calls clip.tokenize(text) to convert the input string into tokens, then passes these tokens to clip.encode_from_tokens_scheduled to generate the conditioning tensor that diffusion samplers use to guide image generation.
What is the purpose of the dynamicPrompts flag in ComfyUI?
The dynamicPrompts flag, defined in comfy_api/latest/_io.py, marks a text input as containing placeholders that require runtime evaluation. When set to True, the execution engine creates a DynamicPrompt wrapper that expands variables like ${artist} using the extra_data dictionary before the node receives the final string.
How do I pass variables for dynamic prompt substitution via the API?
Include an extra_data object at the top level of your JSON payload with key-value pairs matching your placeholder names. For example, "extra_data": {"artist": "Van Gogh"} allows the engine to replace ${artist} with "Van Gogh" before tokenization occurs.
Can I use conditional logic in dynamic prompts?
Yes, the ComfyUI prompt parser supports conditional syntax such as ${if:condition|true text|false text} and random selection with ${choose|option1|option2}. These expressions evaluate during the execution phase managed by comfy_execution/graph.py, producing different prompt variations per workflow run.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →