Asynchronous Video Job Processing in Grok: Architecture and Implementation

Grok handles video generation asynchronously through a job queue pattern where clients submit requests via POST /v1/videos/generations, poll for completion via GET /v1/videos/{id}, and retrieve finalized content via GET /v1/videos/{id}/content, while background workers manage the actual rendering pipeline.

The chenyme/grok2api repository implements this asynchronous architecture to manage long-running video generation tasks without blocking client connections. By decoupling the API interface from the upstream rendering service, the system maintains responsive endpoints while processing computationally intensive video creation. This guide examines the complete flow from job creation to content delivery using the actual source implementation.

How Asynchronous Video Processing Works

The pipeline follows a standard asynchronous job pattern with three distinct phases: job creation, background processing, and status polling.

Job Creation and Queueing

When a client submits a video generation request, the system immediately creates a job record and returns a tracking identifier. In backend/internal/transport/http/inference/handler.go, the generateVideo handler parses the incoming videoGenerationRequest and instantiates a mediadomain.Job struct with StatusInProgress.

The job persistence layer in backend/internal/domain/media/job.go defines the core data structure:

type Job struct {
    ID          string
    Status      mediadomain.Status
    UpstreamURL string
    Seconds     int
    // ... additional metadata
}

This immediate response pattern prevents client timeouts while the actual video rendering occurs in the background.

Background Processing with Providers

Once the job is persisted, control passes to the provider implementation in backend/internal/infra/provider/web/video.go. The provider constructs a videoCreatePayload and communicates with the upstream Grok video service.

The provider handles the entire lifecycle:

  • Transmits the generation request to upstream services
  • Monitors the rendering process
  • Updates the Job record with StatusCompleted, the UpstreamURL, and duration metadata upon completion
  • Handles error states and retries as necessary

This background processing ensures the API remains available while video generation—which may take seconds or minutes—completes.

Polling and Status Retrieval

Clients monitor job progress through dedicated status endpoints. The system translates internal status constants (StatusInProgress, StatusCompleted) to user-friendly string values ("pending", "done", "failed") in JSON responses.

API Endpoints and Handler Implementation

The HTTP layer in backend/internal/transport/http/inference/handler.go exposes three primary routes for video operations:

POST /v1/videos/generations

This endpoint initiates the asynchronous workflow. The handler validates the request body, creates the job record, and returns an immediate acknowledgement.

POST /v1/videos/generations HTTP/1.1
Content-Type: application/json

{
  "model": "grok-imagine-video",
  "prompt": "A sunrise over a futuristic city",
  "duration": "8",
  "image": { "url": "https://example.com/scene.png" }
}

Response:

{
  "request_id": "video_request_123",
  "status": "pending",
  "progress": 0,
  "model": "grok-imagine-video"
}

GET /v1/videos/{request_id}

The getVideo handler retrieves the current job state from the repository and returns a videoGenerationResponse. While processing, the response includes a progress percentage field. Upon completion, the response includes the video object.

Pending status:

{
  "request_id": "video_request_123",
  "status": "pending",
  "progress": 42,
  "model": "grok-imagine-video"
}

Completed status:

{
  "request_id": "video_request_123",
  "status": "done",
  "model": "grok-imagine-video",
  "video": {
    "url": "https://assets.grok.com/video.mp4",
    "duration": 8,
    "respect_moderation": true
  }
}

GET /v1/videos/{request_id}/content

The getVideoContent handler streams the final video file. It utilizes the videoContentURL helper method, which respects hot-updated runtime settings through SetPublicAPIBaseURLResolver. If the job has an UpstreamURL, the API redirects to that location; otherwise, it serves the stored asset directly with appropriate ETag and caching headers.

GET /v1/videos/video_request_123/content HTTP/1.1

Data Models and Persistence Layer

The mediadomain.Job struct in backend/internal/domain/media/job.go serves as the single source of truth for video generation state. Key fields include:

  • ID: Unique identifier correlating requests across the distributed system
  • Status: Current lifecycle state (in-progress, completed, or failed)
  • UpstreamURL: Direct link to the generated asset provided by the Grok video service
  • Seconds: Duration of the generated video content

The repository pattern abstracts persistence details, allowing the handler and provider to interact with job state through domain methods rather than direct database operations.

Response Formats and Content Delivery

The videoGenerationResponse function constructs standardized JSON payloads that differentiate between pending and completed states. For incomplete jobs, the response omits the video field and optionally includes progress. Completed jobs return a fully populated video object containing:

  • url: Accessible endpoint for the video file
  • duration: Length in seconds
  • respect_moderation: Boolean flag indicating content policy compliance

The videoContentURL method dynamically constructs content URLs based on runtime configuration, ensuring that changes to the public API base URL take effect without requiring application restarts.

Summary

  • Asynchronous architecture prevents client timeouts by decoupling API requests from long-running video rendering tasks
  • Three-phase workflow includes job creation (POST /v1/videos/generations), status polling (GET /v1/videos/{id}), and content retrieval (GET /v1/videos/{id}/content)
  • Core implementation resides in backend/internal/transport/http/inference/handler.go with routes handled by generateVideo, getVideo, and getVideoContent
  • State management uses the mediadomain.Job struct defined in backend/internal/domain/media/job.go to track progress from StatusInProgress to StatusCompleted
  • Provider integration in backend/internal/infra/provider/web/video.go handles upstream communication via videoCreatePayload and related orchestration methods
  • Dynamic URL resolution through videoContentURL supports hot-updated runtime configuration via SetPublicAPIBaseURLResolver

Frequently Asked Questions

How does Grok handle long-running video generation without blocking API requests?

The system implements an asynchronous job queue pattern. When a client posts to POST /v1/videos/generations, the generateVideo handler immediately creates a Job record with StatusInProgress and returns a request_id. The actual rendering occurs in background workers within the provider implementation, allowing the API to accept new requests while processing continues.

While the source code does not enforce a specific polling rate, clients should implement exponential backoff when calling GET /v1/videos/{request_id}. Typical implementations poll every 1-2 seconds initially, backing off to 5-10 seconds as processing continues. The response includes a progress field (when available) to indicate completion percentage.

How does the system determine the final video URL returned to clients?

The videoContentURL method in the handler constructs the content URL using SetPublicAPIBaseURLResolver to respect hot-updated configuration. If the Job record contains an UpstreamURL from the provider in backend/internal/infra/provider/web/video.go, the system returns that direct link; otherwise, it generates a URL pointing to the local content endpoint at /v1/videos/{request_id}/content.

What happens if the upstream video generation service fails?

If the background provider process encounters an error, it updates the Job status to the failed state. Subsequent calls to GET /v1/videos/{request_id} return "status": "failed" instead of "done". The error handling preserves the job record for debugging while allowing clients to implement retry logic or error messaging based on the definitive status response.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →