GPT-Image2 Prompt Template Structure: The Complete Guide to Plain-Text and JSON Formats

A standard GPT-Image2 prompt template uses a dual-format architecture with a human-readable plain-text layer and a deterministic JSON schema built around five core fields: type, platform/product/layout, style, content, and constraints.

The awesome-gpt-image-2 repository by freestylefly defines this dual-format system to balance flexibility for manual prompting with precision for automated generation. Whether you are prototyping in a web UI or orchestrating batch image generation through an agent, understanding this template structure is essential for consistent, high-quality outputs.

Plain-Text vs. JSON: The Two-Layer Architecture

The GPT-Image2 template system operates on two interchangeable layers:

Layer Use Case Key Characteristics
Plain-text prompt Direct UI/CLI input, rapid prototyping Free-form description with action verbs, visual specifications, and output constraints
JSON prompt Programmatic generation, agent automation Fixed schema with nested objects for deterministic parsing by downstream APIs

Both layers convey identical information. The plain-text version prioritizes human readability, while the JSON version prioritizes determinism — fixed field names that map directly to model tokens, eliminating ambiguity about which visual elements to render.

The Five Core Fields of the JSON Schema

Every GPT-Image2 prompt template, regardless of category (UI, infographic, poster, illustration), follows an identical top-level structure defined in docs/templates.md. These five fields appear in consistent order:

  1. type — High-level category designation (e.g., "UI Screenshot", "Infographic", "Poster")
  2. platform / product / layout — Concrete specifications for the target context
  3. style — Nested object containing theme, primary_color, typography, and related visual properties
  4. content — Nested object describing structural elements: headings, cards, modules, data fields
  5. constraints — Explicit rules governing aspect ratio, text readability, fidelity level, or prohibited elements

This uniformity allows the gpt-image-2-style-library skill and other automation tools to switch between image categories without redefining the entire schema.

Plain-Text Prompt Example (UI Screenshot)

When working directly in a prompt interface, you can use natural language following this pattern:

为移动健康应用生成一张 iOS 界面图。
核心功能:运动追踪、卡路里统计、社交分享。
视觉风格:极简,主色蓝色,强调色橙色。
布局:顶部导航 + 卡片流,信息层级清晰,留白充足。
输出:高保真 UI 截图,文字清晰可读,比例 9:16。

The plain-text format includes:

  • Action verb initiating the request (为…生成…)
  • Target platform and scene context
  • Visual style, colors, and layout principles
  • Output specifications and constraints

JSON Prompt Example: The Canonical Structure

The same information expressed as structured JSON, as defined in docs/templates.md#tpl-ui:

{
  "type": "UI Screenshot",
  "platform": "iOS",
  "product": "Fitness App",
  "layout": "Card-based feed with bottom tab bar",
  "style": {
    "theme": "Dark Mode",
    "primary_color": "Neon Green",
    "typography": "Clean sans-serif"
  },
  "content": {
    "header": "Today's Activity",
    "cards": [
      { "title": "Running", "data": "5.2 km", "button": "Start" },
      { "title": "Calories", "data": "340 kcal" }
    ]
  },
  "constraints": "High fidelity, readable text, 9:16 aspect ratio"
}

The JSON preserves field order as shown above and can be submitted directly to the generation API or consumed by automation skills.

Using the Template Programmatically

The gpt-image-2-style-library skill exposes this JSON structure for Node.js applications. From src/apimartClient.js, the client wrapper forwards completed prompts to the APIMart generation API:

import { generateImage } from '@freestylefly/gpt-image-2-style-library';

const uiPrompt = {
  type: 'UI Screenshot',
  platform: 'iOS',
  product: 'Fitness App',
  layout: 'Card-based feed with bottom tab bar',
  style: {
    theme: 'Dark Mode',
    primary_color: 'Neon Green',
    typography: 'Clean sans-serif'
  },
  content: {
    header: "Today's Activity",
    cards: [
      { title: 'Running', data: '5.2 km', button: 'Start' },
      { title: 'Calories', data: '340 kcal' }
    ]
  },
  constraints: 'High fidelity, readable text, 9:16 aspect ratio'
};

generateImage(uiPrompt).then(url => console.log('Result URL:', url));

The generateImage helper validates the prompt against the standard schema before transmission, ensuring all required fields are present.

Category-Specific Extensions

While the core five-field structure remains constant, each category adds domain-specific keys:

Category Additional Fields Location in Repository
UI Screenshot platform, layout, detailed content.cards array docs/templates.md#tpl-ui
Infographic data_source, chart_types, narrative_flow docs/templates.md#tpl-infographic
Poster campaign_goal, call_to_action, print_size docs/templates.md#tpl-poster
Illustration art_direction, mood_board_refs, character_specs docs/templates.md#tpl-illustration

The data/style-library.json file centralizes these category definitions, exposing style names, default values, and valid enumeration options that both the web frontend (src/main.jsx) and agent skills consume.

Key Files Defining the Template Architecture

Understanding the GPT-Image2 prompt template requires familiarity with these source files from the freestylefly/awesome-gpt-image-2 repository:

Summary

  • GPT-Image2 prompt templates use a dual-format architecture: plain-text for humans, JSON for machines.
  • The canonical JSON schema contains five ordered fields: type, platform/product/layout descriptors, style, content, and constraints.
  • Field names are fixed across all categories, enabling deterministic API parsing and agent automation without schema redefinition.
  • Category-specific extensions add domain keys while preserving the core structure, as documented in docs/templates.md.
  • The gpt-image-2-style-library skill and apimartClient.js wrapper provide programmatic interfaces for template-based generation.

Frequently Asked Questions

How do I convert a plain-text prompt to the JSON format?

Map each sentence in your plain-text description to the corresponding JSON field: category statements become type, platform references populate platform/product, visual descriptions move into style, structural elements populate content, and output rules become constraints. The data/style-library.json file contains valid values for enumeration fields like theme and primary_color.

Can I use the JSON template with custom fields not in the standard schema?

The standard schema in docs/templates.md defines required fields for deterministic generation. While the API may accept additional keys, agent skills like gpt-image-2-style-library validate against the canonical structure. For experimental fields, include them within the constraints string or extend the schema in a fork, recognizing that downstream tools may ignore unrecognized keys.

Where does the style library get its default values?

Default values and valid enumerations for style fields originate from data/style-library.json, which serves as the single source of truth for both the web interface (src/main.jsx) and programmatic clients. This file is version-controlled and updated when new visual styles or platform targets are added to the GPT-Image2 ecosystem.

What is the difference between layout and style in the JSON template?

The layout field describes spatial organization — structural patterns like "Card-based feed with bottom tab bar" or "Hero section with three-column grid." The style field describes visual treatment — aesthetic properties including theme (Dark Mode/Light Mode), primary_color, typography, and color temperature. Layout answers "how elements are arranged"; style answers "how they appear."

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →