How to Optimize Prompts for Character Consistency in Batch Image Generation

You achieve airtight character consistency across batch generations by treating prompts as structured JSON objects with reusable atomic blocks, then submitting them through a deterministic template engine that concatenates these blocks in a fixed order.

The Awesome-GPT-Image-2 repository implements a Prompt-as-Code architecture that treats prompts as composable data structures rather than free-form text. This approach allows you to define a character’s visual identity once in a canonical JSON block and reference it across dozens or hundreds of generation cases without introducing drift.

Core Architecture

The consistency mechanism relies on five interconnected layers that enforce a pure data contract through JSON.

Atomic Schema

In docs/templates.md, the repository defines an atomic schema that decomposes prompts into discrete, reusable fields such as character, lighting, material, layout, and visual_detail. Each field exists as a plain JSON property that can be concatenated programmatically, ensuring that the character object remains identical across every prompt assembly.

Template Engine

The template engine provides ready-made protocols—including character-portrait, character-grid, and character-storyboard—that embed the character block into larger layouts. As documented in docs/templates.md, the character-grid template specifically demonstrates how the same character description is reused for multiple cells within a single generated image.

Style-Library Skill

Located at agents/skills/gpt-image-2-style-library/SKILL.md, this Vercel-compatible skill loads the library of atomic pieces and exposes the renderTemplate function. It registers the style library and provides a character-type API that agents or scripts can invoke to retrieve canonical character definitions.

Batch API Client

The submitPlatformGeneration function in src/apimartClient.js serves as a thin wrapper around APIMart/HiAPI services. It constructs a payload containing caseId, prompt (the fully-rendered string), and language, then submits these payloads in rapid succession. The function signature preserves the exact prompt text, ensuring the character block remains unmodified during transmission.

The data/cases.json file stores every concrete case, including a dedicated "character" field. Because multiple cases can reference the same character object verbatim, the database acts as a single source of truth for visual identity across batch jobs.

Three-Step Implementation Workflow

To enforce character consistency in batch generation, follow this deterministic pipeline:

  1. Define a canonical character block once in data/style-library.json or a custom file, specifying immutable attributes like facial features, clothing, and color palette.

  2. Reference that block in every case for your batch job—either by copying the JSON directly or importing it via the style-library skill—while varying only peripheral fields like background or lighting.

  3. Generate the batch using the API client (submitPlatformGeneration), which stringifies the entire prompt object. Because the character portion never changes between requests, the server can cache model-conditioning data for that visual identity, improving both speed and cost efficiency.

Practical Code Examples

Defining a Reusable Character Block (JSON)

Store your canonical character definition in data/style-library.json:

{
  "character": {
    "description": "young woman with voluminous dark wavy hair, round wire-rimmed glasses, soft porcelain skin, wearing a sleek navy blazer and white shirt",
    "style": "photorealistic, 8K, soft lighting, shallow depth of field",
    "pose": "standing, hands on hips, slight smile"
  }
}

This block includes only attributes that must remain constant, leaving variable elements like lighting conditions to the surrounding template.

Building Batch Prompts with the Template Engine

Use renderTemplate from the style-library skill and submitPlatformGeneration from the API client to process variations:

import { renderTemplate } from './agents/skills/gpt-image-2-style-library/SKILL.js';
import { submitPlatformGeneration } from './src/apimartClient.js';

const sharedCharacter = require('./data/style-library.json').character;

const variations = [
  { caseId: 101, background: "sunset city skyline", lighting: "golden hour" },
  { caseId: 102, background: "rainy alley", lighting: "neon glow" },
  { caseId: 103, background: "forest clearing", lighting: "dappled sunlight" }
];

async function generateBatch() {
  for (const v of variations) {
    const promptObject = {
      ...sharedCharacter,
      background: v.background,
      lighting: v.lighting
    };
    
    const fullPrompt = renderTemplate('character_grid', promptObject);
    
    await submitPlatformGeneration(
      { caseId: v.caseId, prompt: fullPrompt, language: 'en' },
      fetch
    );
  }
}

generateBatch();

The submitPlatformGeneration function builds the JSON body { caseId, prompt, language } and posts it to the APIMart endpoint, as implemented in src/apimartClient.js.

Using the Character Grid Protocol Directly

For grid layouts that display the same character in multiple cells:

const prompt = renderTemplate('character_grid', {
  character: sharedCharacter,
  rows: 3,
  cols: 2,
  background: "plain white"
});

await submitPlatformGeneration(
  { caseId: 200, prompt, language: 'en' }
);

The character_grid protocol automatically generates a 3×2 layout where the identical character description appears in each cell, guaranteeing perfect visual consistency across the grid.

Why This Approach Works for Batch Generation

Deterministic concatenation: The template engine joins the character block with other fields in a fixed order, eliminating random wording variations that could alter the character's appearance.

Single source of truth: Editing the canonical character block in data/style-library.json automatically propagates changes to every case referencing it, preventing visual drift across large datasets.

Scalable payload structure: Because the character portion remains constant, the generation service can cache the model-conditioning state for that specific visual identity, reducing latency and computational cost for subsequent batch submissions.

Summary

  • Prompt-as-Code architecture transforms character descriptions into reusable JSON blocks that maintain consistency through deterministic assembly.
  • The atomic schema in docs/templates.md separates immutable character traits from variable environmental parameters.
  • Template protocols like character_grid enforce layout consistency while preserving character identity across multiple cells or images.
  • Batch submission via submitPlatformGeneration in src/apimartClient.js transmits the exact same character string in every payload, enabling server-side caching and cost optimization.
  • Centralized storage in data/cases.json ensures that character definitions remain synchronized across thousands of generation cases.

Frequently Asked Questions

What is the Prompt-as-Code approach in Awesome-GPT-Image-2?

Prompt-as-Code treats prompts as structured JSON objects with typed fields rather than unstructured text. According to the repository's docs/templates.md, this allows you to compose prompts programmatically by concatenating atomic blocks—such as character, lighting, and background—ensuring that specific elements like character appearance remain identical across every generation.

How does the character block prevent visual drift across batch jobs?

The character block acts as a single source of truth stored in JSON files like data/style-library.json. When you reference this block across multiple cases in data/cases.json or import it via the style-library skill (agents/skills/gpt-image-2-style-library/SKILL.md), you guarantee that the exact same descriptive string reaches the API every time. The submitPlatformGeneration function in src/apimartClient.js then transmits this string without modification, preventing the subtle prompt variations that typically cause visual inconsistency.

Can I change lighting and background without altering the character's appearance?

Yes. The atomic schema explicitly separates character from environmental fields like lighting and background. By defining these as distinct JSON properties, you can vary the surrounding scene while keeping the character object constant. The template engine concatenates these blocks in a fixed order, ensuring that the character's core description remains untouched even as you iterate through dozens of environmental variations in a batch loop.

Where is the batch generation API logic implemented?

The batch submission logic resides in src/apimartClient.js, specifically within the submitPlatformGeneration function. This function accepts a payload object containing caseId, prompt, and language, then posts it to the APIMart endpoint. For batch workflows, you wrap this function in a loop that iterates over your case variations while maintaining a constant character block, as demonstrated in the repository's implementation examples.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →