Understanding the 6-Block Prompt Construction Pipeline in awesome‑gpt‑image‑2 Agent Skills

The 6-block prompt construction pipeline in awesome‑gpt‑image‑2 is a structured workflow that assembles six logical prompt blocks—Subject & Task, Composition & Layout, Visual Style & Materials, Text & Label Requirements, Aspect Ratio & Output Format, and Constraints & Negative Details—into a single GPT‑Image2 prompt string, as defined in agents/skills/gpt-image-2-style-library/SKILL.md (lines 28‑34).

This repository provides an open‑source agent skill that transforms vague user requests into detailed, production‑ready prompts for OpenAI’s GPT‑Image2 model. By enforcing a strict six‑block structure, the gpt-image-2-style-library skill ensures every generated prompt contains complete visual instructions while maintaining consistency across the style library.

What Is the 6‑Block Prompt Construction Pipeline?

The pipeline is a deterministic assembly process implemented in the gpt-image-2-style-library agent skill. According to lines 22‑27 of agents/skills/gpt-image-2-style-library/SKILL.md, the workflow first detects the input language, determines the target output type, matches a template, and then builds the final prompt by concatenating six predefined logical blocks. This guarantees that every prompt includes mandatory details like subject context, layout rules, material specifications, text requirements, technical parameters, and negative constraints.

Breaking Down the Six Logical Blocks

Each block serves a distinct purpose in the final prompt string. The skill definition maps each block to a specific line in SKILL.md.

1. Subject & Task

Defined in SKILL.md at line 28, this block captures the primary object or concept to be generated and the core creative directive. It answers what should appear in the image and what action the model should perform.

Example content: “Create a Cannes‑level premium summer beverage campaign poster for a fictional lemon drink brand called ‘LIMORA’”.

2. Composition & Layout

Specified at line 29, this block describes the visual arrangement, framing, or grid structure. It controls spatial organization such as column layouts, horizon lines, or subject positioning.

Example content: “Strict 2‑column by 3‑row grid layout with six perfectly aligned panels; thin white dividers; consistent horizon line”.

3. Visual Style & Materials

Located at line 30, this block defines artistic style, textures, lighting scenarios, and material cues. It establishes the aesthetic quality and photorealistic parameters.

Example content: “Ultra‑realistic young woman on a bright sandy beach; oversized lemons; golden‑hour lighting; high‑end advertising aesthetic”.

4. Text & Label Requirements

Documented at line 31, this block declares any textual elements that must appear in the image, including brand names, captions, or UI labels. It ensures the model handles typography correctly.

Example content: “Include the brand name ‘LIMORA’ in each panel; subtle typographic title at the top”.

5. Aspect Ratio & Output Format

Found at line 32, this block sets technical specifications including dimensions, orientation, resolution, and file format.

Example content: “Vertical 4:5, 8K resolution, PNG”.

6. Constraints & Negative Details

Defined across lines 33‑34, this final block lists prohibitions and negative prompts to suppress unwanted artifacts. It acts as a guardrail against common generation errors.

Example content: “No text beyond brand name, avoid UI elements, no watermarks, no extra characters”.

How the Pipeline Works in Practice

The skill’s workflow (lines 22‑27) orchestrates the assembly process. After language detection and template matching, the agent populates each of the six blocks with context‑specific values. The pipeline then concatenates these blocks using a pipe delimiter (|) to produce a clean, parseable string ready for the GPT‑Image2 API.

This structured approach ensures that prompts are complete (no missing technical parameters), consistent (follows the repository’s style library), and portable (the pipe delimiter makes blocks visually distinct for debugging).

Code Example: Building a Production‑Ready GPT‑Image2 Prompt

Below is a runnable JavaScript implementation that mirrors how the awesome‑gpt‑image‑2 agent constructs prompts using the six‑block pipeline.

// Example: Premium beverage poster generation
const promptBlocks = {
  // 1️⃣ Subject & Task
  subjectTask: "Create a Cannes-level premium summer beverage campaign poster for a fictional lemon drink brand called 'LIMORA'",
  
  // 2️⃣ Composition & Layout
  composition: "strict 2-column by 3-row grid layout with six perfectly aligned panels; thin white dividers; consistent horizon line",
  
  // 3️⃣ Visual Style & Materials
  visualStyle: "ultra-realistic young woman on a bright sandy beach; oversized lemons; golden-hour lighting; high-end advertising aesthetic",
  
  // 4️⃣ Text & Label Requirements
  textLabels: "include the brand name 'LIMORA' in each panel; subtle typographic title at the top",
  
  // 5️⃣ Aspect Ratio & Output Format
  output: "vertical 4:5, 8K resolution, PNG",
  
  // 6️⃣ Constraints & Negative Details
  constraints: "no text beyond brand name, avoid UI elements, no watermarks, no extra characters"
};

// Assembly logic: join blocks with pipe delimiter
const finalPrompt = [
  promptBlocks.subjectTask,
  promptBlocks.composition,
  promptBlocks.visualStyle,
  promptBlocks.textLabels,
  promptBlocks.output,
  promptBlocks.constraints
].join(" | ");

console.log(finalPrompt);

Output prompt ready for GPT‑Image2:


Create a Cannes-level premium summer beverage campaign poster for a fictional lemon drink brand called 'LIMORA' | strict 2-column by 3-row grid layout with six perfectly aligned panels; thin white dividers; consistent horizon line | ultra-realistic young woman on a bright sandy beach; oversized lemons; golden-hour lighting; high-end advertising aesthetic | include the brand name 'LIMORA' in each panel; subtle typographic title at the top | vertical 4:5, 8K resolution, PNG | no text beyond brand name, avoid UI elements, no watermarks, no extra characters

Key Source Files and Architecture

The 6-block prompt construction pipeline relies on four critical files that define templates, reference data, and build automation:

Summary

  • The 6-block prompt construction pipeline assembles prompts in a fixed sequence: Subject & Task, Composition & Layout, Visual Style & Materials, Text & Label Requirements, Aspect Ratio & Output Format, and Constraints & Negative Details.
  • Each block is defined in agents/skills/gpt-image-2-style-library/SKILL.md (lines 28‑34) and corresponds to a specific visual or technical requirement.
  • The workflow (lines 22‑27) detects language, matches templates, and concatenates blocks using a pipe (|) delimiter.
  • The pipeline ensures GPT‑Image2 prompts are complete, consistent, and aligned with the repository’s data/style-library.json definitions.

Frequently Asked Questions

How does the 6-block pipeline handle multilingual inputs?

The pipeline includes an initial language detection step (referenced in SKILL.md lines 22‑23) that identifies the source language before template matching. This allows the agent to select appropriate reference materials from data/style-library.json while maintaining the six-block structure regardless of input language.

What delimiter separates the blocks in the final prompt?

The agent joins the six blocks using a pipe character (|) as a delimiter. This creates a visually distinct separation between logical sections, making the prompt easier to debug and parse while ensuring GPT‑Image2 receives a single continuous string.

Can I modify the block order or add custom blocks?

The current implementation in SKILL.md enforces a fixed six-block sequence to maintain compatibility with the style library validation logic. While the generate-style-skill.mjs build script regenerates the skill from JSON sources, modifying the block structure would require updating the schema in data/style-library.json and rebuilding the skill assets.

Where are the negative constraints (block 6) actually processed?

Block 6 (Constraints & Negative Details) is defined at lines 33‑34 of SKILL.md and contains explicit instructions for the model to avoid specific artifacts. These negative prompts are concatenated into the final string like other blocks, but they function as guardrails that suppress unwanted elements such as watermarks, UI chrome, or blurry edges in the generated image.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →