Runtime Modes for the gpt-image-2 Skill: A Complete Technical Guide

The gpt-image-2 skill operates in three distinct runtime modes—Garden local generation, Host-Native delegation, and Advisor—automatically selected by the scripts/check-mode.js detector based on environment variables and host agent capabilities.

The gpt-image-2 skill in the ConardLi/garden-skills repository dynamically adapts its execution strategy to match your runtime environment. According to the skill manifest in skills/gpt-image-2/SKILL.md, the mode detection logic evaluates the presence of the ENABLE_GARDEN_IMAGEGEN and OPENAI_API_KEY environment variables alongside host capabilities to determine whether to generate images locally, delegate to external tools, or function as a prompt-only advisor.

Understanding the Three Runtime Modes

The skill implements a tiered execution model defined in skills/gpt-image-2/scripts/check-mode.js. Each mode represents a different integration strategy between the Garden framework and the underlying image generation infrastructure.

Mode A – Garden Local Generation

Mode A activates when ENABLE_GARDEN_IMAGEGEN is truthy and OPENAI_API_KEY is present in the environment. In this configuration, the skill functions as the primary image generator.

The execution flow follows the full end-to-end pipeline:

  1. Selects an appropriate template from the skill's configuration
  2. Renders the structured prompt
  3. Invokes either scripts/generate.js for text-to-image creation or scripts/edit.js for image modification
  4. Persists the prompt to garden-gpt-image-2/prompt/ and the resulting image to garden-gpt-image-2/image/

This mode requires full API access and local storage permissions, making it suitable for standalone Garden deployments with direct OpenAI integration.

Mode B – Host-Native Delegation

Mode B triggers when ENABLE_GARDEN_IMAGEGEN is unset or false and the host agent exposes its own image generation capabilities. The skill detects host tools such as image_generation, dalle, or nano_banana through the capability negotiation protocol.

In this delegation mode:

  • The skill does not invoke scripts/generate.js or scripts/edit.js
  • It produces only a structured prompt and saves it locally under garden-gpt-image-2/prompt/
  • The rendered prompt is handed off to the host's native image tool
  • The host determines where the final image is stored

This mode is ideal for integrated development environments where the AI assistant provides its own image generation infrastructure, allowing the skill to focus on prompt engineering while leveraging existing tooling.

Mode C – Advisor (Prompt-Only)

Mode C serves as a high-quality prompt consultant when ENABLE_GARDEN_IMAGEGEN is disabled and the host agent lacks image generation capabilities.

The behavior in this mode includes:

  • Rendering optimized prompts and saving them to garden-gpt-image-2/prompt/
  • Printing the prompt to the user interface
  • Recommending external tools such as ChatGPT, Midjourney, or DALL·E for actual image generation
  • No image production occurs locally or through delegation

This fallback mode ensures the skill remains useful even in restricted environments by providing professional-grade prompt engineering that users can transfer to their preferred image generation platform.

How Mode Detection Works

The detection logic resides exclusively in skills/gpt-image-2/scripts/check-mode.js. This lightweight utility evaluates the execution context and returns either a human-readable summary or structured JSON output for programmatic consumption.

To check which mode your environment supports:


# Human-readable output

node skills/gpt-image-2/scripts/check-mode.js

# Structured JSON for automation pipelines

node skills/gpt-image-2/scripts/check-mode.js --json

The script imports shared utilities from skills/gpt-image-2/scripts/shared.js for environment loading and API configuration validation, ensuring consistent behavior across all three modes.

Practical Usage Examples by Mode

Depending on the detected mode, your interaction with the skill changes significantly.

Mode A – Local generation:


# Generate a new image

node skills/gpt-image-2/scripts/generate.js \
  --prompt "A futuristic city skyline at sunset" \
  --size 1024x1024 \
  --quality high

# Edit an existing image

node skills/gpt-image-2/scripts/edit.js \
  --image assets/source.png \
  --prompt "Replace the background with a clean studio scene"

Mode B – Host delegation:


# The skill saves the prompt locally

# Then invoke the host's tool with the rendered prompt

host_image_tool --prompt "$(cat garden-gpt-image-2/prompt/my-task-20260831-101500.md)"

Mode C – Advisor output:

After executing the skill in Mode C, the system displays:

已生成可直接复用的高质量 prompt,请将其粘贴到 ChatGPT / Midjourney 等图像工具中。

Configuration and Environment Variables

The runtime mode selection depends on two critical environment variables defined in the skill's configuration:

  • ENABLE_GARDEN_IMAGEGEN: Must be set to a truthy value (e.g., true, 1) to enable local generation capabilities
  • OPENAI_API_KEY: Required for Mode A to authenticate with OpenAI's image generation APIs

When ENABLE_GARDEN_IMAGEGEN is absent or falsy, the skill inspects the host agent's tool registry to determine whether Mode B (delegation) or Mode C (advisor) is appropriate. This logic ensures zero-configuration deployment in diverse hosting environments while maintaining full functionality when local generation is explicitly enabled.

Summary

  • The gpt-image-2 skill automatically selects from three runtime modes based on environment variables and host capabilities detected by scripts/check-mode.js.
  • Mode A provides full local image generation using scripts/generate.js and scripts/edit.js when ENABLE_GARDEN_IMAGEGEN and OPENAI_API_KEY are configured.
  • Mode B delegates prompt execution to host-native tools like dalle or image_generation when local generation is disabled but host capabilities exist.
  • Mode C operates as a prompt-only advisor, storing outputs in garden-gpt-image-2/prompt/ and recommending external tools when no generation capability is available.
  • All modes preserve prompt history in the local garden-gpt-image-2/prompt/ directory for reuse and version control.

Frequently Asked Questions

How does the gpt-image-2 skill decide which runtime mode to use?

The skill executes skills/gpt-image-2/scripts/check-mode.js at initialization to evaluate the environment. It checks for the ENABLE_GARDEN_IMAGEGEN and OPENAI_API_KEY variables first; if both are present, it selects Mode A. If local generation is disabled, it inspects the host agent's available tools to choose between Mode B (delegation) or Mode C (advisor).

What environment variables are required to activate Mode A?

Mode A requires two conditions: ENABLE_GARDEN_IMAGEGEN must be set to a truthy value, and OPENAI_API_KEY must contain a valid API key. Without both variables, the skill automatically falls back to Mode B or C depending on host capabilities.

Can I manually force the skill to use a specific mode?

According to the source code in skills/gpt-image-2/SKILL.md, the mode is determined automatically by the detection script. While you can influence the selection by setting or unsetting ENABLE_GARDEN_IMAGEGEN, there is no direct manual override parameter—the design emphasizes environment-driven configuration for consistent CI/CD and deployment behavior.

Where are generated images and prompts stored in each mode?

In Mode A, prompts save to garden-gpt-image-2/prompt/ and images to garden-gpt-image-2/image/. In Mode B, only prompts are stored locally in garden-gpt-image-2/prompt/ while images are handled by the host tool. In Mode C, prompts are saved to garden-gpt-image-2/prompt/ and displayed to the user, but no images are produced or stored by the skill.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →