Prompt Building Blocks for Subject, Composition, and Style in Awesome-GPT-Image-2

Awesome-GPT-Image-2 treats every visual prompt as five atomic, composable blocks—Subject, Composition, Style, Lighting/Materials, and Constraints—that agents and authors can mix, match, and serialize into precise natural-language prompts for granular image control.

The Awesome-GPT-Image-2 repository implements a modular schema that splits complex image generation requests into discrete, reusable components. By decomposing prompts into standardized building blocks defined in README.md and docs/templates.md, the system enables both manual crafting and automated agent-driven prompt assembly.

The Five Atomic Building Blocks

The project vision explicitly defines an atomic schema that separates concerns into five logical parts. According to the Vision section in README.md (lines 62-65), these blocks are subjects, lighting, materials, layout, and visual details.

Subject (The "What")

The Subject block defines the primary object, character, product, or scene the model must render. In docs/templates.md, this appears as placeholders like [subject], subjectDescription, or JSON keys such as "subject": "…". This block answers what occupies the visual center of the image.

Composition and Layout (The "How")

The Composition block controls spatial arrangement, aspect ratio, perspective, and hierarchical structure. Template files reference this through fields like layout, composition, structure, and aspect ratio. For example, UI templates specify "Card-based feed with bottom tab bar" while infographic templates define "3×3 grid" or "poster layout" structures.

Style and Aesthetic (The "Look")

The Style block establishes the visual language, including art style, era, mood, color palette, and texture. Defined in docs/templates.md under "Style" sections, this block uses keys like style, theme, art_style, and visuals.style. The data/style-library.json catalog provides controlled vocabulary entries such as "Apple-inspired," "cinematic," or "water-ink" to populate this block consistently.

Lighting and Materials (The "Feel")

The Lighting and Materials blocks dictate how light interacts with surfaces and what tactile qualities the subject exhibits. Photo templates in docs/templates.md specify lighting, materials, and environment.lighting, while camera specifications control physical rendering properties. Architecture templates use this block to define surface textures and environmental illumination.

Constraints and Details (The "Rules")

The Constraints block enforces hard limits such as aspect ratio, resolution, forbidden elements, and quality thresholds. Throughout template files, this appears as constraints, avoid, output, and quality_constraints fields. This block ensures the final output adheres to technical requirements like "9:16 aspect ratio" or "avoid generic stock photography."

How the Blocks Work Together

These building blocks are designed to be independent and composable. A UI template can be reused with different subjects simply by swapping the Subject block, while maintaining the same Layout and Style definitions. As implemented in agents/skills/gpt-image-2-style-library/SKILL.md, agents programmatically assemble complete prompts by filling these blocks in a JSON structure, which the repository then serializes into natural-language instructions for the image model.

This modular architecture means you can switch a prompt from "portrait" to "product" photography by changing the subject and layout fields while preserving the style and lighting configurations.

Practical Implementation

JSON Structure Example

The docs/templates.md file (lines 24-45) demonstrates how these blocks manifest in a structured UI screenshot template:

{
  "type": "UI Screenshot",
  "platform": "iOS",
  "product": "Fitness App",
  "layout": "Card-based feed with bottom tab bar",
  "style": {
    "theme": "Dark Mode",
    "primary_color": "Neon Green",
    "typography": "Clean sans-serif"
  },
  "content": {
    "header": "Today's Activity",
    "cards": [
      {"title": "Running", "data": "5.2 km", "button": "Start"},
      {"title": "Calories", "data": "340 kcal"}
    ]
  },
  "constraints": "High fidelity, readable text, 9:16 aspect ratio"
}

Natural Language Serialization

When serialized for direct prompt input, the blocks translate to specific descriptive clauses:

生成一张[平台,如 iOS]界面图,主题为[极简/科技/拟物]风格,主色[Neon Green],布局为[卡片流 + 底部标签栏],输出高保真 UI 截图,文字清晰可读,比例 9:16。

Advanced Composition Example

For complex creative work, you can explicitly reference each block:

Design a poster for an event titled "[主题]" where the Subject is a stylized rocket launch, the Composition follows a 3×3 grid with the rocket centred, the Style is vintage sci-fi with teal-orange palette, Lighting is dramatic rim-light, and Constraints forbid any modern logos.

Summary

  • Awesome-GPT-Image-2 structures prompts as five atomic blocks defined in the repository Vision: Subject, Composition/Layout, Style, Lighting/Materials, and Constraints.
  • Each block maps to specific JSON keys and template placeholders in docs/templates.md, enabling precise control over image generation parameters.
  • The modular design supports both manual prompt crafting and automated assembly by agents using the gpt-image-2-style-library skill.
  • Blocks are independent—changing the Subject or Layout does not require altering Style or Lighting configurations.
  • Source definitions reside in README.md (lines 62-65) for the conceptual schema and docs/templates.md (lines 24-45) for implementation templates.

Frequently Asked Questions

What are the five building blocks in Awesome-GPT-Image-2?

According to the Vision section in README.md, the five atomic building blocks are Subject (主体), Lighting (光线), Materials (材质), Layout (布局), and Visual Details (视觉细节). These correspond to the Subject, Lighting/Materials, Composition, and Style/Constraints blocks used in practice templates.

How do I change the composition without affecting the style?

Because the schema treats Composition (layout, structure, aspect ratio) and Style (aesthetic, palette, era) as separate blocks, you can modify the layout or composition field in your JSON template or prompt while keeping the style object unchanged. This independence is core to the repository's modular architecture.

Where are the style keywords defined?

The controlled vocabulary for the Style block resides in data/style-library.json, which catalogs terms like "Apple-inspired," "cinematic lighting," and "water-ink." Agents reference this file to ensure consistent style terminology when assembling prompts programmatically.

Can agents automatically assemble these blocks?

Yes. The agents/skills/gpt-image-2-style-library/SKILL.md implementation demonstrates how agents can programmatically fill the five building blocks in a JSON structure and serialize them into natural-language prompts, enabling automated prompt generation while maintaining the atomic schema's precision.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →