Runtime Modes for the gpt-image-2 Skill: A Complete Technical Guide
The gpt-image-2 skill operates in three distinct runtime modes—Garden local generation, Host-Native delegation, and Advisor—automatically selected by the scripts/check-mode.js detector based on environment variables and host agent capabilities.
The gpt-image-2 skill in the ConardLi/garden-skills repository dynamically adapts its execution strategy to match your runtime environment. According to the skill manifest in skills/gpt-image-2/SKILL.md, the mode detection logic evaluates the presence of the ENABLE_GARDEN_IMAGEGEN and OPENAI_API_KEY environment variables alongside host capabilities to determine whether to generate images locally, delegate to external tools, or function as a prompt-only advisor.
Understanding the Three Runtime Modes
The skill implements a tiered execution model defined in skills/gpt-image-2/scripts/check-mode.js. Each mode represents a different integration strategy between the Garden framework and the underlying image generation infrastructure.
Mode A – Garden Local Generation
Mode A activates when ENABLE_GARDEN_IMAGEGEN is truthy and OPENAI_API_KEY is present in the environment. In this configuration, the skill functions as the primary image generator.
The execution flow follows the full end-to-end pipeline:
- Selects an appropriate template from the skill's configuration
- Renders the structured prompt
- Invokes either
scripts/generate.jsfor text-to-image creation orscripts/edit.jsfor image modification - Persists the prompt to
garden-gpt-image-2/prompt/and the resulting image togarden-gpt-image-2/image/
This mode requires full API access and local storage permissions, making it suitable for standalone Garden deployments with direct OpenAI integration.
Mode B – Host-Native Delegation
Mode B triggers when ENABLE_GARDEN_IMAGEGEN is unset or false and the host agent exposes its own image generation capabilities. The skill detects host tools such as image_generation, dalle, or nano_banana through the capability negotiation protocol.
In this delegation mode:
- The skill does not invoke
scripts/generate.jsorscripts/edit.js - It produces only a structured prompt and saves it locally under
garden-gpt-image-2/prompt/ - The rendered prompt is handed off to the host's native image tool
- The host determines where the final image is stored
This mode is ideal for integrated development environments where the AI assistant provides its own image generation infrastructure, allowing the skill to focus on prompt engineering while leveraging existing tooling.
Mode C – Advisor (Prompt-Only)
Mode C serves as a high-quality prompt consultant when ENABLE_GARDEN_IMAGEGEN is disabled and the host agent lacks image generation capabilities.
The behavior in this mode includes:
- Rendering optimized prompts and saving them to
garden-gpt-image-2/prompt/ - Printing the prompt to the user interface
- Recommending external tools such as ChatGPT, Midjourney, or DALL·E for actual image generation
- No image production occurs locally or through delegation
This fallback mode ensures the skill remains useful even in restricted environments by providing professional-grade prompt engineering that users can transfer to their preferred image generation platform.
How Mode Detection Works
The detection logic resides exclusively in skills/gpt-image-2/scripts/check-mode.js. This lightweight utility evaluates the execution context and returns either a human-readable summary or structured JSON output for programmatic consumption.
To check which mode your environment supports:
# Human-readable output
node skills/gpt-image-2/scripts/check-mode.js
# Structured JSON for automation pipelines
node skills/gpt-image-2/scripts/check-mode.js --json
The script imports shared utilities from skills/gpt-image-2/scripts/shared.js for environment loading and API configuration validation, ensuring consistent behavior across all three modes.
Practical Usage Examples by Mode
Depending on the detected mode, your interaction with the skill changes significantly.
Mode A – Local generation:
# Generate a new image
node skills/gpt-image-2/scripts/generate.js \
--prompt "A futuristic city skyline at sunset" \
--size 1024x1024 \
--quality high
# Edit an existing image
node skills/gpt-image-2/scripts/edit.js \
--image assets/source.png \
--prompt "Replace the background with a clean studio scene"
Mode B – Host delegation:
# The skill saves the prompt locally
# Then invoke the host's tool with the rendered prompt
host_image_tool --prompt "$(cat garden-gpt-image-2/prompt/my-task-20260831-101500.md)"
Mode C – Advisor output:
After executing the skill in Mode C, the system displays:
已生成可直接复用的高质量 prompt,请将其粘贴到 ChatGPT / Midjourney 等图像工具中。
Configuration and Environment Variables
The runtime mode selection depends on two critical environment variables defined in the skill's configuration:
ENABLE_GARDEN_IMAGEGEN: Must be set to a truthy value (e.g.,true,1) to enable local generation capabilitiesOPENAI_API_KEY: Required for Mode A to authenticate with OpenAI's image generation APIs
When ENABLE_GARDEN_IMAGEGEN is absent or falsy, the skill inspects the host agent's tool registry to determine whether Mode B (delegation) or Mode C (advisor) is appropriate. This logic ensures zero-configuration deployment in diverse hosting environments while maintaining full functionality when local generation is explicitly enabled.
Summary
- The gpt-image-2 skill automatically selects from three runtime modes based on environment variables and host capabilities detected by
scripts/check-mode.js. - Mode A provides full local image generation using
scripts/generate.jsandscripts/edit.jswhenENABLE_GARDEN_IMAGEGENandOPENAI_API_KEYare configured. - Mode B delegates prompt execution to host-native tools like
dalleorimage_generationwhen local generation is disabled but host capabilities exist. - Mode C operates as a prompt-only advisor, storing outputs in
garden-gpt-image-2/prompt/and recommending external tools when no generation capability is available. - All modes preserve prompt history in the local
garden-gpt-image-2/prompt/directory for reuse and version control.
Frequently Asked Questions
How does the gpt-image-2 skill decide which runtime mode to use?
The skill executes skills/gpt-image-2/scripts/check-mode.js at initialization to evaluate the environment. It checks for the ENABLE_GARDEN_IMAGEGEN and OPENAI_API_KEY variables first; if both are present, it selects Mode A. If local generation is disabled, it inspects the host agent's available tools to choose between Mode B (delegation) or Mode C (advisor).
What environment variables are required to activate Mode A?
Mode A requires two conditions: ENABLE_GARDEN_IMAGEGEN must be set to a truthy value, and OPENAI_API_KEY must contain a valid API key. Without both variables, the skill automatically falls back to Mode B or C depending on host capabilities.
Can I manually force the skill to use a specific mode?
According to the source code in skills/gpt-image-2/SKILL.md, the mode is determined automatically by the detection script. While you can influence the selection by setting or unsetting ENABLE_GARDEN_IMAGEGEN, there is no direct manual override parameter—the design emphasizes environment-driven configuration for consistent CI/CD and deployment behavior.
Where are generated images and prompts stored in each mode?
In Mode A, prompts save to garden-gpt-image-2/prompt/ and images to garden-gpt-image-2/image/. In Mode B, only prompts are stored locally in garden-gpt-image-2/prompt/ while images are handled by the host tool. In Mode C, prompts are saved to garden-gpt-image-2/prompt/ and displayed to the user, but no images are produced or stored by the skill.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →