How the Agent Skill in Awesome-GPT-Image-2 Automates Prompt Construction

The Agent skill automates prompt construction by executing a deterministic 8-step workflow that transforms vague user intent into production-ready GPT-Image-2 prompts through hierarchical template matching, six-block assembly logic, and data-driven constraint enforcement.

The awesome-gpt-image-2 repository provides a specialized Agent-compatible skill called gpt-image-2-style-library that eliminates manual prompt engineering for OpenAI's image models. According to the source code in agents/skills/gpt-image-2-style-library/SKILL.md, this skill systematizes automation by leveraging a structured JSON index to construct, refine, and localize outputs based on repository-backed templates rather than memory or hard-coded defaults.

The 8-Step Deterministic Workflow

The skill follows a strict procedural pipeline encoded in SKILL.md that converts raw requests into precise image generation instructions.

Step 1: Language Detection and Response Alignment

The skill detects the language of the incoming request and locks the entire workflow to that language. If the user inputs Chinese, the Agent processes, reasons, and outputs in Chinese unless explicitly instructed otherwise.

Step 2: Target Output Type Classification

Next, the Agent identifies the intended target output type from a fixed taxonomy of twelve categories: product, poster, UI, infographic, brand, photo, illustration, character, scene, history, document, or special task. This classification filters the available template pool.

Step 3: Hierarchical Template Matching

The skill queries the style-library index (data/style-library.json) using a prioritized matching sequence:

  1. Template category
  2. Visual-style tag
  3. Scene tag
  4. Nearest example cases

This hierarchy ensures the Agent selects the most specific template available before falling back to broader matches.

Step 4: Candidate Presentation and Selection

When multiple templates satisfy the matching criteria, the skill returns 2-3 strong candidates with brief rationales explaining why each fits the request. The skill pauses for explicit user confirmation before proceeding, preventing arbitrary template selection.

Step 5: The Six Building Blocks of Prompt Construction

Upon template selection, the skill assembles the final prompt using six mandatory building blocks:

  • Subject & task – core content and action description
  • Composition & layout – spatial arrangement and structural rules
  • Visual style & materials – aesthetic direction and texture specifications
  • Text & label requirements – typography, language, and readability constraints
  • Aspect ratio & output format – technical specifications such as 16:9 or PNG
  • Constraints & negative details – specific elements to exclude or avoid

Step 6: Copy-Ready Output Generation

The skill formats the finalized prompt for immediate copying, placing it at the top of the response. Metadata follows, including the selected template name and relevant example-case IDs for traceability and reproducibility.

Step 7: Concrete Constraint Enforcement

Before finalizing, the skill validates that all constraints remain concrete and actionable. This includes verifying exact wording for labels, confirming aspect ratio compatibility, ensuring text hierarchy readability, and explicitly listing artifacts to avoid.

Step 8: Final Localization

The skill performs a final localization pass, ensuring Chinese requests produce Chinese prompts (and vice versa) unless the user explicitly requests English output, maintaining linguistic consistency across all prompt components.

Architecture and Data Sources

The automation relies on two primary components that separate data from logic.

The Style Library Index (data/style-library.json)

This JSON file serves as the master data source containing all available templates, categories, visual-style tags, scene tags, known pitfalls, and example cases. The Agent queries this index during the matching phase to ensure selections reflect the most current repository state without embedding data directly into the workflow logic.

The Skill Manifest (SKILL.md)

Located at agents/skills/gpt-image-2-style-library/SKILL.md, this human-readable manifest encodes the deterministic workflow, decision trees, and assembly instructions that the Agent follows. It functions as the execution blueprint, instructing the Agent how to process requests and construct prompts without requiring hard-coded implementations.

Installation and Agent Integration

Install the skill into local Agent environments (Codex, Claude Code, or shared agents) using the provided CLI script:


# Installs to ~/.codex/skills, ~/.claude/skills, or ~/.agents/skills

npm run install:skill

This command executes agents/skills/gpt-image-2-style-library/bin/install.mjs, which copies the skill package into the appropriate Agent skill directories.

Usage and Skill Regeneration

Once installed, Agents invoke the skill programmatically:

const request = "用 gpt-image-2-style-library 技能生成城市生命系统图谱";

agent.invokeSkill('gpt-image-2-style-library', request).then(result => {
  console.log(result.prompt);   // Copy-ready GPT-Image-2 prompt
  console.log(result.template); // Selected template name
  console.log(result.examples);   // Relevant example-case IDs
});

When the underlying data/style-library.json changes, regenerate the skill to synchronize the manifest with the latest templates:

npm run generate:style-skill

This command updates SKILL.md and related references in agents/skills/gpt-image-2-style-library/references/style-library.md to reflect new additions to the style library index.

Summary

  • The Agent skill automates prompt construction through an 8-step deterministic workflow defined in SKILL.md.
  • Hierarchical matching prioritizes template category, then visual-style tags, scene tags, and example cases from data/style-library.json.
  • Prompt assembly uses six building blocks: subject/task, composition/layout, style/materials, text/labels, aspect ratio/format, and constraints/negative details.
  • The skill supports 12 target output types including product, poster, UI, infographic, and illustration.
  • Localization ensures input and output languages match unless explicitly overridden.
  • Installation occurs via npm run install:skill, with regeneration available through npm run generate:style-skill.

Frequently Asked Questions

How does the skill handle ambiguous requests?

When a request matches multiple templates, the skill presents 2-3 strong candidates with brief rationales and asks the user to select one. This prevents arbitrary selection and ensures the final prompt aligns with specific user intent rather than defaulting to the first match.

What happens when the style library is updated?

Run npm run generate:style-skill to regenerate the skill manifest from the latest data/style-library.json. This updates the template pool, tags, and example cases available to the Agent without requiring manual edits to SKILL.md.

Can the skill generate prompts in languages other than English?

Yes. The skill detects the input language at the start of the workflow and localizes the entire process, producing Chinese prompts for Chinese requests (and vice versa) unless the user explicitly requests English output.

Where does the skill store its template data?

Template definitions, categories, and example cases reside in data/style-library.json at the repository root. The skill manifest at agents/skills/gpt-image-2-style-library/SKILL.md contains the procedural logic, while agents/skills/gpt-image-2-style-library/references/style-library.md provides a human-readable reference derived from the JSON index.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →