How to Implement Text Encoding with CLIP and Dynamic Prompts in ComfyUI

To implement text encoding with CLIP and dynamic prompts in ComfyUI, load a CLIP model using the CLIPLoader node, then pass the model and a text string to the CLIPTextEncode node, ensuring the text input is marked with dynamicPrompts=True to enable runtime placeholder expansion.

ComfyUI provides a modular architecture for Stable Diffusion workflows where text encoding with CLIP and dynamic prompts enables runtime prompt modification without changing the node graph structure. This guide examines the source code implementation in the Comfy-Org/ComfyUI repository to show how the CLIPTextEncode node processes text with CLIP models and how the execution engine resolves placeholders before tokenization.

CLIP Text Encoding Architecture

The implementation involves three core components that work together to transform text into conditioning tensors:

  • CLIPLoader (nodes.py lines 975-1002): Loads a CLIP or CLIP-Vision checkpoint and makes the model available to downstream nodes via the CLIP output type.

  • CLIPTextEncode (nodes.py lines 59-81): Accepts a CLIP model and a text string, tokenizes the input using clip.tokenize(text), and returns a conditioning tensor via clip.encode_from_tokens_scheduled that samplers consume.

  • DynamicPrompt Engine (comfy_execution/graph.py lines 22-38): Wraps the prompt dictionary and expands placeholders like ${artist} before the node executes, storing the original prompt and an ephemeral copy for substitutions.

The execution flow follows this sequence: First, CLIPLoader produces a CLIP model instance. Next, CLIPTextEncode receives this model along with a text input. If the text input declares "dynamicPrompts": True in its definition (as seen in comfy_api/latest/_io.py lines 25-34), the engine creates a DynamicPrompt wrapper when the workflow submits (as implemented in execution.py lines 704-712). The wrapper expands any placeholders using variables from the extra_data dictionary, then the node tokenizes the final string.

How Dynamic Prompts Work

The dynamic prompt system operates through three distinct layers that handle declaration, serialization, and runtime expansion.

Widget Declaration

In the node definition, the text input specifies dynamicPrompts=True within its input options. This flag, defined in comfy_api/latest/_io.py, instructs the frontend that the field may contain Jinja-style expressions or placeholder variables.

Runtime Expansion

The DynamicPrompt class in comfy_execution/graph.py maintains two internal dictionaries:

self.original_prompt   # Raw JSON received from the client

self.ephemeral_prompt  # Mutable copy where dynamic text is substituted

When execution.py processes a workflow (lines 704-712), it instantiates this wrapper. As the topological sorter executes each node, the engine evaluates placeholders against the extra_data dictionary passed in the API request or built-in helpers like ${prompt}.

Evaluation Logic

The expansion logic resides in comfy/prompt_parser.py (invoked by the DynamicPrompt wrapper). This parser resolves syntax such as ${variable}, ${choose|option1|option2}, and ${if:condition|true text|false text} before the text ever reaches the CLIP tokenizer.

Dynamic Prompt Usage Patterns

Depending on your workflow requirements, you can implement several patterns for text encoding:

  • Static prompts: Provide a plain string such as "a portrait of a cyberpunk city" with no special syntax required.

  • Variable substitution: Pass variables in the extra_data payload and reference them as "A painting by ${artist}". The engine substitutes values before tokenization.

  • Random selection: Use the built-in ${choose|option1|option2|option3} syntax to randomly select alternatives during each execution.

  • Conditional blocks: Implement branching logic with ${if:condition|true text|false text} to modify prompts based on runtime conditions.

All evaluations occur once per execution, and the resulting final string feeds into CLIPTextEncode.

Implementation Examples

Static Prompt Workflow

The minimal JSON workflow demonstrates a basic text encoding without dynamic elements:

{
  "1": {
    "class_type": "CLIPLoader",
    "inputs": {
      "ckpt_name": "clip-vit-large-patch14.safetensors"
    }
  },
  "2": {
    "class_type": "CLIPTextEncode",
    "inputs": {
      "clip": ["1", 0],
      "text": "A futuristic city at sunset"
    }
  }
}

The text field contains a literal string that the node tokenizes directly without expansion.

Dynamic Prompt with Placeholders

To substitute variables at runtime, include extra_data in your payload:

{
  "extra_data": {
    "artist": "Studio Ghibli"
  },
  "1": {
    "class_type": "CLIPLoader",
    "inputs": { "ckpt_name": "clip-vit-large-patch14.safetensors" }
  },
  "2": {
    "class_type": "CLIPTextEncode",
    "inputs": {
      "clip": ["1", 0],
      "text": "A beautiful landscape in the style of ${artist}"
    }
  }
}

When executed, the engine replaces ${artist} with Studio Ghibli before the CLIPTextEncode node processes the text.

Python API Submission

You can programmatically submit workflows with dynamic prompts using the ComfyUI API:

import json, urllib.request

# Build the prompt JSON with variables

prompt = {
    "extra_data": {"artist": "Alex Ross"},
    "1": {"class_type": "CLIPLoader",
          "inputs": {"ckpt_name": "clip-vit-large-patch14.safetensors"}},
    "2": {"class_type": "CLIPTextEncode",
          "inputs": {"clip": ["1", 0],
                     "text": "A heroic portrait in the style of ${artist}"}}
}

# POST to the ComfyUI API

data = json.dumps({"prompt": prompt}).encode()
req = urllib.request.Request("http://127.0.0.1:8188/prompt", data=data)
urllib.request.urlopen(req)   # Execution is queued

The request includes extra_data, so the server resolves ${artist} to Alex Ross before the CLIP node receives the final string.

Custom Node with Dynamic Support

To create a custom node that supports dynamic prompts, declare the input with the dynamicPrompts flag:

class MyPromptNode:
    @classmethod
    def INPUT_TYPES(s):
        return {
            "required": {
                "text": (
                    "STRING",
                    {"multiline": True, "dynamicPrompts": True,
                     "tooltip": "Enter a prompt; can contain ${variables}."}
                )
            }
        }

    RETURN_TYPES = ("STRING",)
    FUNCTION = "process"

    def process(self, text):
        # text is already the expanded string

        return (text,)

Because the input definition mirrors CLIPTextEncode, the engine automatically expands placeholders before calling the process method.

Key Source Files

Understanding these specific source files helps when debugging or extending the functionality:

  • nodes.py (lines 59-81): Contains the CLIPTextEncode class that tokenizes and encodes text with CLIP models.

  • nodes.py (lines 975-1002): Implements CLIPLoader, which supplies the CLIP model used by text encoding nodes.

  • comfy_api/latest/_io.py (lines 25-34): Defines the String widget type and the dynamic_prompts flag that serializes into the JSON payload.

  • comfy_execution/graph.py (lines 22-38): Houses the DynamicPrompt wrapper class that manages prompt expansion at runtime.

  • execution.py (lines 704-712): Handles the creation of DynamicPrompt objects for incoming workflow requests.

  • script_examples/basic_api_example.py (lines 53-71): Demonstrates client-side workflow construction and API submission patterns.

Summary

Implementing text encoding with CLIP and dynamic prompts in ComfyUI requires understanding the interaction between model loading, node execution, and runtime text expansion:

  • Load CLIP models using CLIPLoader from nodes.py before any text encoding operations.
  • Use CLIPTextEncode to tokenize text and generate conditioning tensors for samplers.
  • Enable dynamic prompts by setting dynamicPrompts=True in the input definition, allowing the engine to resolve ${variables} before tokenization.
  • Pass substitution variables via the extra_data field in API requests to modify prompts at runtime without changing the workflow graph.
  • Reference comfy_execution/graph.py and execution.py to understand how the DynamicPrompt wrapper manages prompt expansion.

Frequently Asked Questions

How does the CLIPTextEncode node tokenize input text?

The CLIPTextEncode node in nodes.py calls clip.tokenize(text) to convert the input string into tokens, then passes these tokens to clip.encode_from_tokens_scheduled to generate the conditioning tensor that diffusion samplers use to guide image generation.

What is the purpose of the dynamicPrompts flag in ComfyUI?

The dynamicPrompts flag, defined in comfy_api/latest/_io.py, marks a text input as containing placeholders that require runtime evaluation. When set to True, the execution engine creates a DynamicPrompt wrapper that expands variables like ${artist} using the extra_data dictionary before the node receives the final string.

How do I pass variables for dynamic prompt substitution via the API?

Include an extra_data object at the top level of your JSON payload with key-value pairs matching your placeholder names. For example, "extra_data": {"artist": "Van Gogh"} allows the engine to replace ${artist} with "Van Gogh" before tokenization occurs.

Can I use conditional logic in dynamic prompts?

Yes, the ComfyUI prompt parser supports conditional syntax such as ${if:condition|true text|false text} and random selection with ${choose|option1|option2}. These expressions evaluate during the execution phase managed by comfy_execution/graph.py, producing different prompt variations per workflow run.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →