How to Generate Structured JSON Prompts for GPT-Image2 Agents

Structured JSON prompts treat image generation instructions as executable code rather than ambiguous text, enabling programmatic control over style, layout, and constraints.

The GPT-Image2 ecosystem in the freestylefly/awesome-gpt-image-2 repository defines prompts as JSON objects that agents consume directly. This architecture eliminates the guesswork of natural language descriptions and makes automation pipelines, CI/CD workflows, and bot integrations predictable and reproducible.


Understanding the Three-Layer Architecture

The repository organizes prompt generation into three cooperating layers:

  • Prompt Templates — Domain-specific JSON schemas for every visual category
  • Style Library — A canonical database of 500+ style definitions
  • Agent Skill — An NPM package that assembles complete prompts programmatically

Each layer builds upon the previous, allowing you to work at the abstraction level that matches your use case.


Step 1: Select a Template from docs/templates.md

Templates are organized by visual domain. Each entry provides both a human-readable explanation and a machine-parseable JSON skeleton.

Available categories include:

  • UI & Interfaces
  • Infographic & Information Visualization
  • Posters & Marketing Materials
  • Product Photography
  • Architecture & Interiors
  • Photography Styles
  • Illustration & Art
  • Character Design
  • Scene & Environment
  • Historical Reconstruction
  • Document & Data Visualization

The JSON templates use descriptive placeholder keys like [product], [platform], and [audience] that you replace with concrete values.

UI Screenshot Template Example

{
  "type": "UI Screenshot",
  "platform": "iOS",
  "product": "Fitness App",
  "layout": "Card‑based feed with bottom tab bar",
  "style": {
    "theme": "Dark Mode",
    "primary_color": "Neon Green",
    "typography": "Clean sans‑serif"
  },
  "content": {
    "header": "Today's Activity",
    "cards": [
      { "title": "Running",   "data": "5.2 km", "button": "Start" },
      { "title": "Calories",  "data": "340 kcal" }
    ]
  },
  "constraints": "High fidelity, readable text, 9:16 aspect ratio"
}

Source: docs/templates.md, UI & Interfaces section [templates.md†L24-L46]


Step 2: Apply Style Definitions from data/style-library.json

The data/style-library.json file contains a curated collection of style entries, each with:

  • Unique identifier
  • Descriptive keywords
  • Color palette specifications
  • Layout hints and aesthetic guidance

You can reference styles directly by ID or let the skill match based on your high-level description. The flat JSON structure makes programmatic lookup straightforward in any language.


Step 3: Generate Complete Prompts with the Agent Skill

The gpt-image-2-style-library skill automates template selection, style matching, and JSON assembly. Located at agents/skills/gpt-image-2-style-library/SKILL.md, this NPM package exposes a CLI interface for headless operation.

CLI Usage


# Generate an infographic prompt automatically

npx gpt-image-2-style-library generate \
  --type infographic \
  --topic "Urban Metabolism" \
  --audience "General Public"

The skill performs three operations:

  1. Retrieves the matching template structure from docs/templates.md
  2. Queries data/style-library.json for appropriate style entries
  3. Merges user variables, style data, and constraints into a single valid JSON object

Infographic Output Example

{
  "type": "Infographic",
  "topic": "Urban Metabolism",
  "audience": "General Public",
  "structure": {
    "title_area": "城市生命系统图谱",
    "layout": "Isometric cutaway, 12 numbered panels",
    "modules": [
      { "title": "能源", "icon": "lightning", "text": "Power flows" },
      { "title": "水循环", "icon": "water_drop", "text": "Water flows" }
    ]
  },
  "style": {
    "aesthetic": "Scientific atlas",
    "colors": "Low saturation, color‑coded flows",
    "background": "Light paper texture"
  },
  "constraints": "No cyber‑punk, no gibberish text, strict structural layout"
}

Source: docs/templates.md, Infographic section [templates.md†L7-L29]


JSON Schema Design for Validation

The structured JSON prompts use a consistent top-level schema that enables early validation:

Key Purpose Example Values
type Visual category discriminator "UI Screenshot", "Infographic"
platform / topic / product Subject identifier "iOS", "Urban Metabolism"
layout / structure Spatial composition rules "Card‑based feed", "Isometric cutaway"
style Aesthetic parameters from style library Nested object with colors, typography
content Data payload for image elements Arrays of cards, modules, or fields
constraints Quality and exclusion directives Aspect ratios, "no text garble"

Because every key has explicit semantics, downstream services in api/generate-image.js can reject malformed requests before invoking expensive image generation APIs like APIMart or HiAPI.


Deploying to Generation Endpoints

The final JSON payload routes through api/generate-image.js, which validates the structure and forwards to configured backends. The endpoint accepts the complete prompt object and handles provider-specific authentication and rate limiting.

To integrate into automation pipelines, structure your generation workflow as:

1. Define variables (product, topic, audience)
2. Call skill CLI or construct JSON manually
3. POST to /api/generate-image.js with validated payload
4. Handle response (image URL or error details)

Summary

  • JSON prompts are code: The GPT-Image2 ecosystem treats generation instructions as structured data, not free text
  • Templates in docs/templates.md provide domain-specific schemas for every visual category
  • data/style-library.json supplies 500+ canonical style definitions for consistent aesthetics
  • gpt-image-2-style-library skill automates assembly: install via NPM and call the CLI for headless operation
  • Flat, self-describing schema enables validation at api/generate-image.js before expensive API calls

Frequently Asked Questions

How do I extend the JSON prompt schema for custom use cases?

Modify the template structure in docs/templates.md following the existing key naming conventions. Add your custom keys under content or introduce new top-level fields alongside constraints. The api/generate-image.js handler passes through unknown keys, so backend providers can implement extensions without breaking existing clients.

Can I use the style library without the NPM skill?

Yes. data/style-library.json is plain JSON—parse it directly in Python, TypeScript, or any language. Query by id or filter on keywords to retrieve matching entries, then manually merge the style object into your prompt under the style key.

What validation occurs at the API endpoint?

api/generate-image.js checks for required top-level keys (type, content) and validates that style contains recognized entries when strict mode is enabled. Missing or malformed fields return 400 errors with descriptive messages, preventing wasted API calls to image generation services.

How do I prevent "text garble" in generated images?

Include explicit constraints in the constraints field: "readable text only", "no gibberish characters", or "verified font rendering". The templates in docs/templates.md demonstrate proven constraint phrasing that correlates with higher text accuracy in GPT-Image2 outputs.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →