How to Use the gpt-image-2 Skill for Image Generation and Editing
The gpt-image-2 skill is a multi-mode prompt-engineered tool that generates or edits images through OpenAI-compatible APIs, adapting its behavior to run locally, delegate to host tools, or act as a prompt advisor based on environment configuration.
The gpt-image-2 skill from the ConardLi/garden-skills repository provides a flexible interface for AI-powered image generation and editing. It operates in three distinct runtime modes—from full local execution to simple prompt engineering—making it adaptable to various development environments. Whether you need to generate UI mockups, product visuals, or perform complex image edits, this skill offers both CLI scripts and advisory capabilities.
Understanding the Three Runtime Modes
The skill automatically detects its operating environment and adjusts functionality accordingly.
Mode A: Garden Local (Full Execution)
When both ENABLE_GARDEN_IMAGEGEN=1 and a valid OPENAI_API_KEY are present, the skill runs in Garden local mode. In this mode, the skill executes CLI scripts directly to call OpenAI-compatible image APIs.
- Generation script:
skills/gpt-image-2/scripts/generate.jshandles text-to-image requests - Editing script:
skills/gpt-image-2/scripts/edit.jsprocesses image modifications - Storage: Prompts save to
garden-gpt-image-2/prompt/and images togarden-gpt-image-2/image/
Mode B: Host-Native (Prompt Engineering Advisor)
If ENABLE_GARDEN_IMAGEGEN is unset but the host agent provides a built-in image tool, the skill acts as a prompt engineering advisor. It renders structured prompts and hands them to the host's image generator without making direct API calls.
Mode C: Advisor Only (Prompt Generation)
When neither Garden nor a host image tool is available, the skill returns the finalized prompt for manual use. It optionally saves the prompt locally but does not attempt any API calls.
Detecting Your Current Mode
Before executing image operations, determine which mode is active by running the mode detection script:
node skills/gpt-image-2/scripts/check-mode.js
Add the --json flag for machine-readable output suitable for downstream automation:
node skills/gpt-image-2/scripts/check-mode.js --json
The script inspects environment variables and host capabilities to emit mode = A, B, or C.
Core Workflow Structure
All modes follow a standardized six-step workflow:
- Detect mode using
check-mode.js - Select a template from
references/(over 80 JSON/markdown templates across 18 categories) - Fill template arguments with user-provided values or defaults
- Render the final prompt as plain text or JSON
- Persist the prompt in
garden-gpt-image-2/prompt/(mandatory for Modes A and C, recommended for B) - Execute the operation based on current mode capabilities
Generating Images with generate.js (Mode A)
The generate.js script serves as the primary CLI entry point for text-to-image generation. It parses command-line flags and constructs JSON payloads for the /images/generations endpoint.
Basic Image Generation
Generate a high-quality image by specifying prompt, size, and quality parameters:
node skills/gpt-image-2/scripts/generate.js \
--prompt "A cute baby sea otter" \
--size 1024x1024 \
--quality high \
--output garden-gpt-image-2/image/otter.png \
--json
The script's buildPayload function (lines 34-46 in generate.js) constructs the API request body containing model, size, and quality parameters. It then uses postJson from shared.js to send the request to https://api.openai.com/v1/images/generations, extracts the base-64 image bytes, and writes the PNG file to the specified output path.
Generation from Saved Prompts
Reuse structured prompts by referencing a saved prompt file:
node skills/gpt-image-2/scripts/generate.js \
--promptfile garden-gpt-image-2/prompt/my-poster.md \
--output garden-gpt-image-2/image/poster.png
Editing Images with edit.js (Mode A)
The edit.js script handles image modification through the /images/edits endpoint, supporting both full-image edits and masked region edits.
Background Replacement
Replace entire backgrounds by providing a source image and edit prompt:
node skills/gpt-image-2/scripts/edit.js \
--image assets/source.png \
--prompt "Replace the background with a clean studio scene" \
--output garden-gpt-image-2/image/edited.png
Local Object Replacement with Masks
For precise edits targeting specific areas, provide a mask image where transparent areas indicate regions to modify:
node skills/gpt-image-2/scripts/edit.js \
--image assets/source.png \
--mask assets/mask.png \
--prompt "Replace only the masked area with a glass vase" \
--output garden-gpt-image-2/image/obj-replaced.png
Host-Native and Advisor-Only Usage (Modes B and C)
When running in Mode B or C, the scripts function as prompt engineering tools without executing API calls.
Mode B: Host-Native Integration
In Mode B, the script outputs structured prompts for consumption by host image tools:
node skills/gpt-image-2/scripts/generate.js \
--prompt "A futuristic UI mockup for a SaaS dashboard" \
--json
The output provides a finalized prompt that you can feed into your host's built-in image_generation tool.
Mode C: Advisor-Only Output
For environments without image capabilities, save the engineered prompt for external use:
node skills/gpt-image-2/scripts/generate.js \
--prompt "A minimalist poster for a music festival" \
--prompt-output garden-gpt-image-2/prompt/festival.md
This saves the prompt to a timestamped file (e.g., festival-20260828-112233.md) and displays a confirmation message with usage instructions.
Implementation Details and Shared Utilities
The skills/gpt-image-2/scripts/shared.js file provides core functionality used across all scripts:
loadAmbientEnv: Reads environment variables from.env,.gateway.env, and~/.gateway.env(lines 32-44)readPromptInput: Handles prompt input from CLI arguments or filessavePrompt: Persists prompts to the designated directorypostJsonandpostMultipart: API callers that automatically attachOPENAI_API_KEYheaders (lines 41-45)- File-system helpers:
ensureFilesExist,slugify, andmakeTimestampfor organized asset management
The skill organizes over 80 prompt templates in the references/ directory, categorized into 18 groups including UI mockups, product visuals, and editing workflows.
Summary
- The gpt-image-2 skill operates in three modes: full local execution (A), host-native advisory (B), and prompt-only advisor (C).
- Mode A requires
ENABLE_GARDEN_IMAGEGEN=1andOPENAI_API_KEY, executinggenerate.jsandedit.jsdirectly against OpenAI-compatible APIs. - Modes B and C focus on prompt engineering, outputting structured prompts for external tools or manual use.
- Core scripts rely on
shared.jsutilities for environment loading, API authentication, and file management. - Generated content saves to
garden-gpt-image-2/prompt/andgarden-gpt-image-2/image/when running locally.
Frequently Asked Questions
How do I determine which mode the gpt-image-2 skill is running in?
Run node skills/gpt-image-2/scripts/check-mode.js to detect the current runtime mode. The script examines the ENABLE_GARDEN_IMAGEGEN environment variable and the presence of OPENAI_API_KEY to determine whether to execute locally (Mode A), advise a host tool (Mode B), or operate as a prompt-only advisor (Mode C).
Can I use the gpt-image-2 skill without an OpenAI API key?
Yes, in Mode C (Advisor only) or Mode B (Host-Native). Without an API key, the skill functions as a prompt engineering tool, saving finalized prompts to garden-gpt-image-2/prompt/ for use with external image generators. However, Mode A requires a valid OPENAI_API_KEY to execute local API calls.
Where does the skill store generated images and prompts?
In Mode A, prompts save to garden-gpt-image-2/prompt/ and generated images to garden-gpt-image-2/image/. The shared.js utilities handle file naming using slugify and makeTimestamp functions to create organized, timestamped files. In Modes B and C, prompt storage is optional based on the --prompt-output flag.
What is the difference between using edit.js with and without a mask?
Without a mask, edit.js treats the entire image as editable, typically used for background replacement or global style changes. With the --mask flag, the script sends a multipart request to /images/edits specifying only the masked (transparent) regions for modification, enabling precise object replacement while preserving the rest of the image. The mask functionality relies on postMultipart in shared.js to handle the multipart/form-data payload.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →