How to Use the gpt-image-2 Skill for Image Generation and Editing in Garden
The gpt-image-2 skill in ConardLi/garden-skills operates in three runtime modes—Garden local, Host-Native, and Advisor-only—to generate or edit images via OpenAI-compatible APIs, using generate.js for text-to-image creation and edit.js for image manipulation.
The gpt-image-2 skill is a multi-mode prompt-engineered tool designed for OpenAI-compatible image generation endpoints. As implemented in the ConardLi/garden-skills repository, it dynamically adapts its execution strategy based on environment configuration and host capabilities, offering over 80 structured prompt templates organized into 18 categories.
Understanding the Three Runtime Modes
The skill detects its operating mode by checking environment variables and host capabilities. According to skills/gpt-image-2/scripts/check-mode.js, the system categorizes execution into three distinct modes:
Mode A – Garden Local activates when ENABLE_GARDEN_IMAGEGEN=1 and a valid OPENAI_API_KEY are present. In this mode, the skill executes CLI scripts directly, saving prompts to garden-gpt-image-2/prompt/ and images to garden-gpt-image-2/image/.
Mode B – Host-Native triggers when ENABLE_GARDEN_IMAGEGEN is unset or false, but the host agent provides a built-in image tool. The skill acts solely as a prompt-engineering advisor, rendering structured prompts for the host's generator without making API calls.
Mode C – Advisor Only engages when neither Garden nor host image tools are available. The skill returns finalized prompts without attempting execution, optionally persisting them locally for external use.
Detecting Your Runtime Mode
Always begin by running the mode detection script to determine current capabilities. The check-mode.js script inspects environment variables and emits the active mode as A, B, or C.
node skills/gpt-image-2/scripts/check-mode.js
For machine-readable output suitable for downstream automation, append the --json flag.
Core Workflow for Image Operations
The gpt-image-2 skill follows a consistent six-step workflow across all modes, implemented in the CLI entry points and shared utilities.
- Detect the mode using
check-mode.js - Select a template from the
references/directory (80+ JSON/markdown templates across 18 categories including UI mockups and product visuals) - Fill template arguments with user-provided values or defaults
- Render the final prompt as plain text or JSON
- Persist the prompt in
garden-gpt-image-2/prompt/(mandatory for Modes A and C, recommended for B) - Execute the operation based on the detected mode
Generating Images with generate.js
For Mode A (Garden local), the skills/gpt-image-2/scripts/generate.js CLI handles text-to-image generation. The script parses flags including --prompt, --size, --quality, --background, and --output, then constructs a JSON payload via the buildPayload function (lines 34-46) for the /images/generations endpoint.
Generate a 1024×1024 high-quality image:
node skills/gpt-image-2/scripts/generate.js \
--prompt "A cute baby sea otter" \
--size 1024x1024 \
--quality high \
--output garden-gpt-image-2/image/otter.png \
--json
The script reads your prompt, constructs a payload containing model, size, and quality parameters, posts JSON to https://api.openai.com/v1/images/generations, extracts base-64 image bytes, and writes the PNG file to the specified output path.
Generate from a saved prompt file:
node skills/gpt-image-2/scripts/generate.js \
--promptfile garden-gpt-image-2/prompt/my-poster.md \
--output garden-gpt-image-2/image/poster.png
Editing Images with edit.js
For image manipulation, skills/gpt-image-2/scripts/edit.js provides CLI access to the /images/edits endpoint. It supports background replacement and masked editing operations.
Replace an image background:
node skills/gpt-image-2/scripts/edit.js \
--image assets/source.png \
--prompt "Replace the background with a clean studio scene" \
--output garden-gpt-image-2/image/edited.png
Edit with a mask for localized object replacement:
node skills/gpt-image-2/scripts/edit.js \
--image assets/source.png \
--mask assets/mask.png \
--prompt "Replace only the masked area with a glass vase" \
--output garden-gpt-image-2/image/obj-replaced.png
Host-Native and Advisor-Only Usage
When operating in Mode B (Host-Native), the skill functions as a prompt engineering layer without direct API access. Execute the scripts normally—they will output the finalized prompt without posting to OpenAI.
Generate structured output for host consumption:
node skills/gpt-image-2/scripts/generate.js \
--prompt "A futuristic UI mockup for a SaaS dashboard" \
--json
Feed this output into your host's built-in image_generation tool.
For Mode C (Advisor-only), save prompts locally without image generation:
node skills/gpt-image-2/scripts/generate.js \
--prompt "A minimalist poster for a music festival" \
--prompt-output garden-gpt-image-2/prompt/festival.md
The script saves the prompt and displays usage instructions for external GPT-Image-2 compatible tools.
Shared Utilities and Implementation Details
The skills/gpt-image-2/scripts/shared.js module provides core functionality used by both generate.js and edit.js.
Environment Loading: The loadAmbientEnv function (lines 32-44) reads configuration from .env, .gateway.env, and ~/.gateway.env files, ensuring OPENAI_API_KEY is available for API calls.
Prompt Handling: Utilities include readPromptInput for reading prompt files and savePrompt for persisting generated prompts to the garden-gpt-image-2/prompt/ directory.
API Communication: The postJson and postMultipart functions automatically attach the OPENAI_API_KEY header when communicating with OpenAI-compatible endpoints.
File System Helpers: Functions like ensureFilesExist, slugify, and makeTimestamp manage directory creation and filename generation for organized asset storage.
Summary
- The gpt-image-2 skill operates in three modes: Garden local (full API execution), Host-Native (prompt advisor), and Advisor-only (prompt generation without API calls)
- Mode detection occurs via
skills/gpt-image-2/scripts/check-mode.js, which checksENABLE_GARDEN_IMAGEGENandOPENAI_API_KEYenvironment variables - Image generation uses
generate.jswith thebuildPayloadfunction to construct requests for the/images/generationsendpoint - Image editing uses
edit.jswith support for masking via the--maskflag and background replacement - Shared utilities in
shared.jshandle environment loading, prompt persistence, and API authentication - The skill includes 80+ prompt templates in the
references/directory covering 18 categories including UI mockups and product visuals
Frequently Asked Questions
How do I determine which mode the gpt-image-2 skill is running in?
Run node skills/gpt-image-2/scripts/check-mode.js to detect the current runtime mode. The script inspects the ENABLE_GARDEN_IMAGEGEN and OPENAI_API_KEY environment variables to determine whether it should execute in Garden local (Mode A), Host-Native (Mode B), or Advisor-only (Mode C).
What is the difference between generate.js and edit.js in the gpt-image-2 skill?
generate.js creates new images from text prompts via the /images/generations endpoint, while edit.js modifies existing images through the /images/edits endpoint. The edit script accepts additional parameters like --mask for localized edits and --image for the source file path.
Where does the gpt-image-2 skill store generated prompts and images?
In Mode A (Garden local), the skill saves prompts to garden-gpt-image-2/prompt/ and generated images to garden-gpt-image-2/image/. The shared.js utility functions handle file persistence using timestamped filenames and directory creation via ensureFilesExist.
Can I use the gpt-image-2 skill without an OpenAI API key?
Yes, in Modes B and C. Without ENABLE_GARDEN_IMAGEGEN=1 or a valid OPENAI_API_KEY, the skill functions as a prompt engineering tool, rendering structured prompts from its 80+ templates in the references/ directory without making API calls. Use --prompt-output to save these prompts for use with external image generators.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →