Grok API Endpoints for Responses, Chat, Messages, Images, and Videos: A Complete Technical Guide

Grok2API exposes OpenAI-compatible endpoints under /v1 for text completions (/v1/responses, /v1/chat/completions), Anthropic Messages (/v1/messages), image generation/editing (/v1/images/*), and asynchronous video jobs (/v1/videos/*).

Grok2API is a Go-based gateway that bridges client applications to Grok's upstream services (Build, Web, and Console). Understanding the available Grok API endpoints is essential for integrating text generation, conversational AI, and media creation into your applications. This guide examines the routing, request handling, and provider architecture implemented in the chenyme/grok2api repository.

HTTP API Surface Overview

All public inference endpoints are mounted under the /v1 path prefix and follow OpenAI or Anthropic API conventions. The route definitions live in backend/internal/transport/http/inference/handler.go, where the Register method maps HTTP verbs to handler functions:

func (h *Handler) Register(router *gin.RouterGroup) {
    router.GET("/models", h.listModels)
    router.POST("/responses", h.createResponse)
    router.POST("/chat/completions", h.createChatCompletion)
    router.POST("/messages", h.createMessage)
    router.POST("/images/generations", h.generateImage)
    router.POST("/images/edits", h.editImage)
    router.POST("/videos/generations", h.generateVideo)
    router.GET("/videos/:requestId", h.getVideo)
    router.GET("/videos/:requestId/content", h.getVideoContent)
    router.POST("/responses/compact", h.compactResponse)
    router.GET("/responses/:responseId", h.getResponse)
    router.DELETE("/responses/:responseId", h.deleteResponse)
}

The Grok API endpoints cover five functional categories:

  • Model Discovery: GET /v1/models lists available models and their routing information.
  • Text Completions: POST /v1/responses and POST /v1/chat/completions handle stateless and conversational text generation.
  • Anthropic Compatibility: POST /v1/messages provides Claude-style request/response formats.
  • Image Operations: POST /v1/images/generations and POST /v1/images/edits create and modify images.
  • Video Operations: POST /v1/videos/generations submits asynchronous jobs, with GET /v1/videos/{id} and GET /v1/videos/{id}/content for polling and retrieval.

Text Generation Endpoints

OpenAI-Style Responses (/v1/responses)

The POST /v1/responses endpoint, handled by createResponse in handler.go, executes stateless text completions. It accepts a JSON payload containing model, input (prompt), and optional stream parameters. When streaming is enabled, the gateway returns Server-Sent Events (SSE) with incremental text deltas.

curl http://127.0.0.1:8000/v1/responses \
  -H "Authorization: Bearer g2a_abcdef12345" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "grok-chat-fast",
        "input": "Explain quantum tunnelling in three sentences.",
        "stream": false
      }'

The gateway routes this request to the appropriate upstream provider (Grok Web by default for grok-chat-fast) and returns a JSON response containing the generated text and metadata.

Chat Completions (/v1/chat/completions)

For conversation-aware interactions, POST /v1/chat/completions (handler: createChatCompletion) accepts a messages array with role-based objects (system, user, assistant). This endpoint supports function calling and streaming via the streamProtocolChat implementation.

fetch('http://localhost:8000/v1/chat/completions', {
  method: 'POST',
  headers: {
    'Authorization': 'Bearer g2a_abcdef12345',
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({
    model: 'grok-chat-fast',
    stream: true,
    messages: [{role: 'user', content: 'Write a haiku about autumn.'}]
  })
}).then(res => {
  const reader = res.body.getReader()
  // Consume the SSE stream and display each delta…
})

Anthropic Messages Compatibility (/v1/messages)

Grok2API provides first-class support for Anthropic's Messages API via POST /v1/messages (handler: createMessage). This endpoint translates Claude-style requests into Grok upstream calls, enabling drop-in compatibility with existing Anthropic SDKs. The request format accepts model, messages, max_tokens, and system parameters, returning a response structure matching Anthropic's schema.

Media Generation Endpoints

Image Generation and Editing

The image endpoints support both creation and modification workflows. POST /v1/images/generations (handler: generateImage) accepts parameters including prompt, size (e.g., "1024x1024"), and response_format ("url" or "b64_json"). The POST /v1/images/edits endpoint (handler: editImage) allows masking and editing existing images via multipart/form-data or JSON payloads with base64-encoded sources.

import requests

url = "http://localhost:8000/v1/images/generations"
headers = {
    "Authorization": "Bearer g2a_abcdef12345",
    "Content-Type": "application/json"
}
payload = {
    "model": "grok-imagine-image-lite",
    "prompt": "A futuristic city at sunset",
    "size": "1024x1024",
    "response_format": "url"
}
resp = requests.post(url, headers=headers, json=payload)
print(resp.json()["data"][0]["url"])

Media size limits are enforced by constants in handler.go (e.g., maxMediaResponseTransferBytes = 2 GB), preventing oversized transfers from upstream providers.

Video Generation (Asynchronous)

Video creation follows an asynchronous job pattern. Submitting a POST request to /v1/videos/generations (handler: generateVideo) returns immediately with a request_id, while the actual rendering occurs in the background.


# 1️⃣ Submit a video job

curl -X POST http://localhost:8000/v1/videos/generations \
  -H "Authorization: Bearer g2a_abcdef12345" \
  -H "Content-Type: application/json" \
  -d '{
        "model": "grok-imagine-video",
        "prompt": "A dragon flying over a medieval castle",
        "duration": "10s"
      }'

# => {"request_id":"vid_12345"}

# 2️⃣ Poll for status

curl http://localhost:8000/v1/videos/vid_12345

# => {"status":"completed"}

# 3️⃣ Download the finished video

curl -OJ http://localhost:8000/v1/videos/vid_12345/content

The job lifecycle is managed by the video generation handler and persisted in the media store (backend/internal/domain/media). Clients poll GET /v1/videos/{requestId} until status indicates completion, then retrieve binary content via GET /v1/videos/{requestId}/content.

Request Flow and Architecture

When calling any Grok API endpoint, requests traverse four architectural layers:

  1. Access Layer: The middleware.ClientKey component in backend/internal/transport/http/middleware/auth.go extracts bearer tokens and enforces rate limits or model allow-lists.

  2. Gateway Routing: The gateway service (backend/internal/application/gateway/service.go) inspects the model prefix (e.g., grok-chat-fast, grok-imagine-image) to select the appropriate provider channel—Build, Web, or Console.

  3. Provider Adapters: Each provider implements the Provider interface defined in backend/internal/infra/provider/definition.go. For example, backend/internal/infra/provider/web/adapter.go handles image and video operations, while backend/internal/infra/provider/build/adapter.go manages text responses.

  4. Egress Management: The egress manager (backend/internal/infra/egress/manager.go) constructs HTTP/SOCKS connections, rotates proxy pools, and handles Cloudflare clearance (FlareSolverr) before forwarding to Grok upstream services.

Key Implementation Files

Component File Path Purpose
HTTP Router backend/internal/transport/http/inference/handler.go Registers all /v1/* routes and defines request structs (imageGenerationRequest, videoGenerationRequest).
Gateway Logic backend/internal/application/gateway/service.go Routes requests to Build, Web, or Console providers based on model metadata.
Provider Interface backend/internal/infra/provider/definition.go Defines operation types (responses, chat, messages, images, videos) and adapter selection logic.
Web Provider backend/internal/infra/provider/web/adapter.go Implements image generation and video job submission for Grok Web.
Build Provider backend/internal/infra/provider/build/adapter.go Handles text completions and streaming for Grok Build.
Egress Manager backend/internal/infra/egress/manager.go Manages outbound connections, proxy fallback, and request retries.
Authentication backend/internal/transport/http/middleware/auth.go Validates client keys and attaches authentication context.

Summary

  • Grok2API provides OpenAI-compatible Grok API endpoints under the /v1 prefix, supporting text, chat, Anthropic Messages, images, and videos.
  • Text endpoints (/v1/responses, /v1/chat/completions) support both single-turn completions and multi-turn conversations with streaming SSE output.
  • Media endpoints follow distinct patterns: images return synchronous base64/URL data, while videos use asynchronous job polling via /v1/videos/{id}.
  • Provider architecture automatically routes requests to Grok Build, Web, or Console based on model prefixes, handled by adapters in backend/internal/infra/provider/.
  • Request flow enforces authentication via middleware, routes through the gateway service, and egresses through managed proxy pools with Cloudflare clearance support.

Frequently Asked Questions

What authentication method do the Grok API endpoints use?

All endpoints require a Bearer token in the Authorization header (e.g., Authorization: Bearer g2a_abcdef12345). The middleware.ClientKey component in backend/internal/transport/http/middleware/auth.go validates these tokens against the configured client key database, enforcing optional model allow-lists and rate limits.

Can I use standard OpenAI client libraries with Grok2API?

Yes. Because Grok2API implements the OpenAI REST schema for /v1/responses, /v1/chat/completions, and /v1/images/generations, you can point standard OpenAI SDKs to your Grok2API base URL. The request and response formats are designed to be drop-in compatible, though some Grok-specific parameters (like model name prefixes) may differ.

How does video generation differ from image generation in the API?

Image generation (/v1/images/generations) is synchronous, returning base64-encoded data or URLs immediately. Video generation (/v1/videos/generations) is asynchronous: it returns a request_id immediately, and you must poll GET /v1/videos/{request_id} to check completion status before downloading content from /v1/videos/{request_id}/content.

Which upstream Grok service handles each type of endpoint?

The gateway routes requests based on model prefixes and operation types. Grok Web typically handles image and video operations (backend/internal/infra/provider/web/adapter.go), while Grok Build manages text responses and chat completions (backend/internal/infra/provider/build/adapter.go). Grok Console provides additional support for stateless responses and media tasks.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →