How to Generate Structured JSON Prompts for GPT-Image2 Agents
Structured JSON prompts treat image generation instructions as executable code rather than ambiguous text, enabling programmatic control over style, layout, and constraints.
The GPT-Image2 ecosystem in the freestylefly/awesome-gpt-image-2 repository defines prompts as JSON objects that agents consume directly. This architecture eliminates the guesswork of natural language descriptions and makes automation pipelines, CI/CD workflows, and bot integrations predictable and reproducible.
Understanding the Three-Layer Architecture
The repository organizes prompt generation into three cooperating layers:
- Prompt Templates — Domain-specific JSON schemas for every visual category
- Style Library — A canonical database of 500+ style definitions
- Agent Skill — An NPM package that assembles complete prompts programmatically
Each layer builds upon the previous, allowing you to work at the abstraction level that matches your use case.
Step 1: Select a Template from docs/templates.md
Templates are organized by visual domain. Each entry provides both a human-readable explanation and a machine-parseable JSON skeleton.
Available categories include:
- UI & Interfaces
- Infographic & Information Visualization
- Posters & Marketing Materials
- Product Photography
- Architecture & Interiors
- Photography Styles
- Illustration & Art
- Character Design
- Scene & Environment
- Historical Reconstruction
- Document & Data Visualization
The JSON templates use descriptive placeholder keys like [product], [platform], and [audience] that you replace with concrete values.
UI Screenshot Template Example
{
"type": "UI Screenshot",
"platform": "iOS",
"product": "Fitness App",
"layout": "Card‑based feed with bottom tab bar",
"style": {
"theme": "Dark Mode",
"primary_color": "Neon Green",
"typography": "Clean sans‑serif"
},
"content": {
"header": "Today's Activity",
"cards": [
{ "title": "Running", "data": "5.2 km", "button": "Start" },
{ "title": "Calories", "data": "340 kcal" }
]
},
"constraints": "High fidelity, readable text, 9:16 aspect ratio"
}
Source: docs/templates.md, UI & Interfaces section [templates.md†L24-L46]
Step 2: Apply Style Definitions from data/style-library.json
The data/style-library.json file contains a curated collection of style entries, each with:
- Unique identifier
- Descriptive keywords
- Color palette specifications
- Layout hints and aesthetic guidance
You can reference styles directly by ID or let the skill match based on your high-level description. The flat JSON structure makes programmatic lookup straightforward in any language.
Step 3: Generate Complete Prompts with the Agent Skill
The gpt-image-2-style-library skill automates template selection, style matching, and JSON assembly. Located at agents/skills/gpt-image-2-style-library/SKILL.md, this NPM package exposes a CLI interface for headless operation.
CLI Usage
# Generate an infographic prompt automatically
npx gpt-image-2-style-library generate \
--type infographic \
--topic "Urban Metabolism" \
--audience "General Public"
The skill performs three operations:
- Retrieves the matching template structure from
docs/templates.md - Queries
data/style-library.jsonfor appropriate style entries - Merges user variables, style data, and constraints into a single valid JSON object
Infographic Output Example
{
"type": "Infographic",
"topic": "Urban Metabolism",
"audience": "General Public",
"structure": {
"title_area": "城市生命系统图谱",
"layout": "Isometric cutaway, 12 numbered panels",
"modules": [
{ "title": "能源", "icon": "lightning", "text": "Power flows" },
{ "title": "水循环", "icon": "water_drop", "text": "Water flows" }
]
},
"style": {
"aesthetic": "Scientific atlas",
"colors": "Low saturation, color‑coded flows",
"background": "Light paper texture"
},
"constraints": "No cyber‑punk, no gibberish text, strict structural layout"
}
Source: docs/templates.md, Infographic section [templates.md†L7-L29]
JSON Schema Design for Validation
The structured JSON prompts use a consistent top-level schema that enables early validation:
| Key | Purpose | Example Values |
|---|---|---|
type |
Visual category discriminator | "UI Screenshot", "Infographic" |
platform / topic / product |
Subject identifier | "iOS", "Urban Metabolism" |
layout / structure |
Spatial composition rules | "Card‑based feed", "Isometric cutaway" |
style |
Aesthetic parameters from style library | Nested object with colors, typography |
content |
Data payload for image elements | Arrays of cards, modules, or fields |
constraints |
Quality and exclusion directives | Aspect ratios, "no text garble" |
Because every key has explicit semantics, downstream services in api/generate-image.js can reject malformed requests before invoking expensive image generation APIs like APIMart or HiAPI.
Deploying to Generation Endpoints
The final JSON payload routes through api/generate-image.js, which validates the structure and forwards to configured backends. The endpoint accepts the complete prompt object and handles provider-specific authentication and rate limiting.
To integrate into automation pipelines, structure your generation workflow as:
1. Define variables (product, topic, audience)
2. Call skill CLI or construct JSON manually
3. POST to /api/generate-image.js with validated payload
4. Handle response (image URL or error details)
Summary
- JSON prompts are code: The GPT-Image2 ecosystem treats generation instructions as structured data, not free text
- Templates in
docs/templates.mdprovide domain-specific schemas for every visual category data/style-library.jsonsupplies 500+ canonical style definitions for consistent aestheticsgpt-image-2-style-libraryskill automates assembly: install via NPM and call the CLI for headless operation- Flat, self-describing schema enables validation at
api/generate-image.jsbefore expensive API calls
Frequently Asked Questions
How do I extend the JSON prompt schema for custom use cases?
Modify the template structure in docs/templates.md following the existing key naming conventions. Add your custom keys under content or introduce new top-level fields alongside constraints. The api/generate-image.js handler passes through unknown keys, so backend providers can implement extensions without breaking existing clients.
Can I use the style library without the NPM skill?
Yes. data/style-library.json is plain JSON—parse it directly in Python, TypeScript, or any language. Query by id or filter on keywords to retrieve matching entries, then manually merge the style object into your prompt under the style key.
What validation occurs at the API endpoint?
api/generate-image.js checks for required top-level keys (type, content) and validates that style contains recognized entries when strict mode is enabled. Missing or malformed fields return 400 errors with descriptive messages, preventing wasted API calls to image generation services.
How do I prevent "text garble" in generated images?
Include explicit constraints in the constraints field: "readable text only", "no gibberish characters", or "verified font rendering". The templates in docs/templates.md demonstrate proven constraint phrasing that correlates with higher text accuracy in GPT-Image2 outputs.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →