How Prompt Building Blocks Are Assembled in the Awesome-GPT-Image-2 Skill Workflow
The gpt-image-2-style-library skill assembles production-ready GPT-Image-2 prompts by deterministically combining six logical building blocks—subject, layout, style, text, format, and constraints—selected from the repository-wide style-library.json based on detected user intent.
The freestylefly/awesome-gpt-image-2 repository implements a "Prompt-as-Code" philosophy that transforms free-form user requests into structured, atomic prompt components. This assembly process ensures every generated prompt follows a consistent schema, eliminating ambiguity while maximizing image generation fidelity.
The Three-Stage Assembly Pipeline
The skill follows a deterministic pipeline defined in agents/skills/gpt-image-2-style-library/SKILL.md. Each stage refines the user's input into increasingly specific prompt components.
Stage 1: Language and Intent Detection
First, the skill analyzes the incoming request to establish foundational parameters. It detects the user's language and infers the output type from a predefined taxonomy that includes UI, poster, infographic, brand, photo, illustration, character, scene, history, document, or special task categories. This classification determines which template branch to query in the subsequent stage.
Stage 2: Template and Tag Matching
Using the detected intent, the skill traverses a hierarchical matching system against data/style-library.json:
- Template category → Visual style tag → Scene tag → Nearest example cases
The JSON file contains exhaustive metadata for every template, including category definitions, style descriptors, scene classifications, associated tags, example case IDs, and instructional arrays. The skill locates the best-fit template definition by matching these hierarchical attributes against the user's implicit requirements.
Stage 3: Prompt Block Construction
Once the skill selects a template, it constructs the final prompt by populating six logical blocks with data from the template's guidance and pitfalls arrays, merged with user-specific inputs. The ordering of these blocks is hard-coded in the SKILL.md workflow (steps 5-6), ensuring consistent atomic structure across all outputs.
The Six Core Building Blocks
The prompt building blocks follow a strict assembly order that mirrors professional prompt engineering best practices:
-
Subject & Task – Defines what the image depicts and its high-level purpose, typically drawn from
tmpl.titleor user input. -
Composition & Layout – Specifies spatial arrangement, visual hierarchy, and aspect-ratio constraints, populated from the template's guidance array.
-
Visual Style & Materials – Applies style tags (e.g., Poster, 3D, UI Screenshot) and material descriptors including lighting and color palette definitions.
-
Text & Label Requirements – Mandates readable text, captions, or labeling rules necessary for the output type.
-
Aspect Ratio & Output Format – declares target dimensions, file types, or platform-specific constraints (defaulting to 16:9 when unspecified).
-
Constraints & Negative Details – Explicit prohibitions sourced from the template's pitfalls array, such as "no watermarks" or "avoid generic backgrounds."
Each block concatenates into a copy-ready string with line breaks separating logical sections, optionally appending the selected template name and example case IDs for reference.
Implementation Example
The following JavaScript pseudo-code illustrates the assembly logic implemented in the skill workflow:
// Pseudo-code: assemblePrompt(request)
function assemblePrompt(request) {
const lang = detectLanguage(request); // Stage 1
const intent = inferIntent(request); // Stage 1
const tmpl = matchTemplate(intent); // Stage 2 (queries style-library.json)
// Stage 3: Build six blocks
const blocks = [
// 1. Subject & Task
`Subject: ${request.subject || tmpl.title.en}`,
// 2. Composition & Layout
`Layout: ${tmpl.guidance.en.join(' ')}`,
// 3. Visual Style & Materials
`Style: ${tmpl.styles.map(s => s.title.en).join(', ')}`,
// 4. Text & Label Requirements
`Text: ${request.text ?? 'Include clear labels'}`,
// 5. Aspect Ratio & Output Format
`Aspect Ratio: ${request.ratio ?? '16:9'}`,
// 6. Constraints & Negative Details
`Constraints: ${tmpl.pitfalls.en.join(' ')}`
];
// Join blocks with line breaks for a copy-ready prompt
return blocks.join('\n');
}
When a user requests "城市生命系统图谱" (City Life System Map), the skill selects the "UI Screenshot System" template (ID ui-screenshot-system) and produces:
Subject: 城市生命系统图谱
Layout: Lock platform, aspect ratio, layout hierarchy, and exact visible text. Specify UI chrome such as status bars, tabs, action rows, or comment layers.
Style: UI
Text: Include clear labels for each subsystem
Aspect Ratio: 16:9
Constraints: Avoid vague platform names and generic app mockups.
Source File Architecture
The assembly mechanism relies on three critical files:
-
agents/skills/gpt-image-2-style-library/SKILL.md– Contains the declarative workflow definition that dictates the six-block assembly order and detection logic. -
data/style-library.json– The master data source enumerating every template, style hierarchy, and associated guidance/pitfalls arrays used to populate the building blocks. -
agents/skills/gpt-image-2-style-library/references/style-library.md– A human-readable markdown view of the JSON data, automatically generated for developer reference.
Summary
-
The freestylefly/awesome-gpt-image-2 skill uses a three-stage pipeline (detection → matching → construction) to ensure deterministic prompt generation.
-
Six building blocks (subject, layout, style, text, format, constraints) are assembled in a hard-coded order defined in SKILL.md.
-
Template selection relies on hierarchical matching against style-library.json, which provides the guidance and pitfalls arrays feeding each block.
-
The "Prompt-as-Code" approach guarantees that outputs follow the atomic schema of subject → layout → style → text → format → constraints, regardless of input language or complexity.
Frequently Asked Questions
What data source feeds the prompt building blocks?
The style-library.json file at data/style-library.json serves as the master data source. It contains exhaustive metadata for every template, including categories, styles, scenes, tags, example case IDs, and the guidance and pitfalls arrays that populate the six building blocks.
How does the skill determine which template category to use?
The skill infers the output type during Stage 1: Language and Intent Detection by analyzing the user's request against a taxonomy that includes UI, poster, infographic, brand, photo, illustration, character, scene, history, document, and special task categories. This classification drives the hierarchical template matching process.
Can developers modify the order of the prompt building blocks?
No. According to the SKILL.md workflow specification (steps 5-6), the ordering of the six blocks is hard-coded to maintain consistency with the "Prompt-as-Code" philosophy. This rigid structure ensures that every generated prompt follows the atomic schema required for optimal GPT-Image-2 performance.
Where is the assembly logic documented for agent developers?
The complete assembly workflow is documented in agents/skills/gpt-image-2-style-library/SKILL.md, which defines the declarative steps for language detection, intent inference, template matching, and the sequential construction of the six prompt building blocks.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →