How to Generate AI Media Using the QuickDesign Plugin: A Complete Guide

The QuickDesign plugin supplies a unified CLI (quickdesign) that wraps hosted AI models—including Seedance 2.0 R2V, Nano Banana 2, and Topaz—for creating AI-generated images and videos through OAuth authentication, model selection, and reference-based workflows.

The quickdesign CLI tool, maintained in the anthropics/claude-plugins-community repository, abstracts complex video generation pipelines into simple command-line operations. By following the canonical workflows defined in the skill definition and pipeline documentation, you can produce high-quality UGC video content with consistent voice continuity and automated post-processing. This guide explains how to generate AI media using the QuickDesign plugin from initial authentication through final publication.

Authentication and Model Registry Setup

Before generating any media, you must authenticate the CLI and verify model availability.

Run the OAuth flow once to store credentials locally:

quickdesign auth login

After authentication, query the live registry to confirm which models are available for your account:


# List available video generation models

quickdesign video models

# List available image generation models

quickdesign image models

According to the SKILL.md file in the repository, the default model for UGC videos is Seedance 2.0 R2V (seedance-2.0-r2v), while the default for image edits is Nano Banana 2 (nano-banana-2). These defaults are optimized for high-quality character continuity and rapid iteration respectively.

Image Generation and Editing Workflows

For image editing tasks—such as changing angles, poses, or object states—use the Nano Banana 2 model with explicit reference image syntax.

Reference Image Syntax

The QuickDesign plugin uses the @ImageN pattern (defined in references/multi-reference-pattern.md) rather than verbose prose to identify visual elements. Pass each reference using the --reference-image flag and reference them in your prompt with @Image1, @Image2, etc.

quickdesign image generate \
  --model nano-banana-2 \
  --reference-image ./avatar.jpg \
  --reference-image ./product.jpg \
  --aspect-ratio 9:16 --resolution 2K \
  -p "@Image1 holds @Image2. Edit @Image1: change pose to selfie. No music score. No subtitles or on-screen text." \
  -o ./edited-avatar.png --wait

Key parameters include:

  • --model: Specifies the provider (e.g., nano-banana-2)
  • --aspect-ratio: Aspect ratio for the output (e.g., 9:16 for mobile vertical)
  • --resolution: Target resolution (e.g., 2K, 1080p)
  • -p: The prompt string using @ImageN references
  • --wait: Blocks until generation completes

Video Generation Strategies

The plugin supports two primary video workflows: single-segment generation for short clips and multi-segment pipelines for longer UGC content with voice continuity.

Single-Segment Generation

For videos under 15 seconds, a single Seedance R2V call is sufficient. This approach requires no intermediate steps or audio extraction:

quickdesign video generate \
  --provider seedance \
  --reference-image ./edited-avatar.png \
  --aspect-ratio 9:16 --duration 12 --resolution 1080p \
  -p '@Image1 in a kitchen. She says: "Try our new smoothie!" No music score. No subtitles or on-screen text.' \
  -o ./segment1.mp4 --wait

The --provider flag selects the backend model family, while --duration accepts values in seconds (typically 4-15 seconds per segment for optimal quality).

Multi-Segment UGC Pipeline

For longer content, follow the canonical UGC pipeline documented in pipelines/ugc-video.md. This workflow maintains voice continuity across segments by extracting audio from the first segment and passing it as a reference to subsequent parallel generations.

Step 1: Generate the first segment sequentially to establish the voice baseline:

quickdesign video generate \
  --provider seedance \
  --reference-image ./avatar.jpg \
  --duration 12 --aspect-ratio 9:16 --resolution 1080p \
  -p '@Image1 greets the viewer. She says: "Welcome!" No music score. No subtitles or on-screen text.' \
  -o seg1.mp4 --wait

Step 2: Extract the audio track for voice cloning:

quickdesign video extract-audio seg1.mp4 -o seg1-audio.mp3

Step 3: Generate remaining segments in parallel, passing the extracted audio with --reference-audio:


# Generate segment 2 in background

quickdesign video generate \
  --provider seedance \
  --reference-image ./avatar.jpg \
  --reference-audio seg1-audio.mp3 \
  --duration 12 --aspect-ratio 9:16 --resolution 1080p \
  -p '@Image1 shows the product. She says: "It tastes amazing!" No music score. No subtitles or on-screen text.' \
  -o seg2.mp4 --wait &

# Generate segment 3 in background

quickdesign video generate \
  --provider seedance \
  --reference-image ./avatar.jpg \
  --reference-audio seg1-audio.mp3 \
  --duration 8 --aspect-ratio 9:16 --resolution 1080p \
  -p '@Image1 gives a call-to-action. She says: "Buy now!" No music score. No subtitles or on-screen text.' \
  -o seg3.mp4 --wait &

wait  # Wait for all parallel jobs to finish

Step 4: Concatenate segments (mux-only, no re-encoding):

quickdesign video concat seg1.mp4 seg2.mp4 seg3.mp4 -o final.mp4

Planning and Cost Estimation

Before executing any paid generation, the QuickDesign plugin requires plan approval. Count words in your script and split content into 4-15 second segments (approximately 30 words per minute). The system generates a plan summary listing the model, duration, resolution, and estimated cost that must be approved before proceeding—even in auto-mode, this plan is displayed for confirmation per the rules in references/confirmation-rules.md.

Post-Processing and Publishing

After generation, apply optional post-processing steps:

Add Subtitles:

quickdesign video subtitle final.mp4 \
  --style tiktok --language en \
  -o final-subbed.mp4 --wait

Upscale Resolution:

quickdesign video upscale final-subbed.mp4 \
  --model topaz-video-upscale \
  -o final-4k.mp4 --wait

Deploy to Meta Ads: When targeting Meta advertising platforms, use the quickdesign meta command family documented in references/deploy-meta.md to push assets directly to your ad account.

Summary

  • Authentication is handled once via quickdesign auth login, storing OAuth credentials locally for subsequent calls.
  • Model selection defaults to Nano Banana 2 for image edits and Seedance 2.0 R2V for UGC videos, with availability confirmed through quickdesign video models.
  • Reference handling requires the @ImageN syntax in prompts and explicit --reference-image flags for visual continuity, plus --reference-audio extracted from Segment 1 for voice continuity in multi-segment workflows.
  • Multi-segment generation follows the canonical pipeline: generate Segment 1, extract audio, generate Segments 2-N in parallel with audio references, then concatenate.
  • Post-processing includes optional subtitle generation via quickdesign video subtitle and upscaling through models like Topaz.

Frequently Asked Questions

How do I maintain the same voice across multiple video segments?

Extract the audio from your first generated segment using quickdesign video extract-audio seg1.mp4 -o audio.mp3, then pass this file to subsequent generation commands with the --reference-audio audio.mp3 flag. This ensures voice continuity as implemented in the pipelines/ugc-video.md workflow.

What is the difference between --provider and --model flags?

The --provider flag selects the backend service family (e.g., seedance), while --model specifies the exact model version (e.g., seedance-2.0-r2v or nano-banana-2). According to the source code, most commands are model-agnostic, and swapping models only requires updating these flags and any model-specific prompt details found in the corresponding model card.

How do I reference multiple images in a single prompt?

Use the @ImageN syntax where N corresponds to the order of your --reference-image arguments. For example, if you pass --reference-image ./avatar.jpg --reference-image ./product.jpg, reference them in your prompt as @Image1 and @Image2 respectively. This pattern is defined in references/multi-reference-pattern.md.

Can I generate video segments in parallel to save time?

Yes, after generating Segment 1 and extracting its audio, you can generate Segments 2 through N in parallel by running the commands in the background (using & in bash) and waiting for completion with the wait command. Each parallel job should include --reference-audio pointing to the extracted audio from Segment 1 to maintain consistency.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →