How to Integrate Grok2API with an Existing Application: Complete Technical Guide

To integrate Grok2API with an existing application, run the Go-based server using environment variables for configuration, obtain JWT tokens via the authentication endpoint, and consume RESTful endpoints for inference, model management, and media handling.

Grok2API is a production-ready Go server that exposes a RESTful HTTP interface for AI inference and model orchestration. According to the chenyme/grok2api source code, the system follows hexagonal (clean) architecture principles organized into distinct transport, application, domain, and infrastructure layers. This guide demonstrates how to integrate Grok2API with an existing application using HTTP clients, Go module imports, or Docker containerization.

Understanding the Grok2API Architecture

Grok2API implements a hexagonal architecture that separates concerns into distinct layers, making integration straightforward regardless of your tech stack.

Architecture Layers

The codebase in backend/internal/ follows these clear boundaries:

  • Transport Layer (transport/http/*): Handles HTTP routing, authentication, and request/response formatting. Key sub-packages include inference, model, media, dashboard, and audit.
  • Application Layer (application/*): Contains business logic and service implementations such as settings, quotarecovery, and model services.
  • Domain Layer (domain/*): Defines pure business entities including model, media, inference, and clientkey.
  • Infrastructure Layer (infra/*): Provides concrete implementations for external concerns including security (JWT), egress (outbound HTTP), runtime (storage), and provider (web/console adapters).

The server bootstrap in backend/cmd/grok2api/main.go wires these components together, creating a dependency injection container that initializes the HTTP server.

Configuration and Deployment Options

Before integrating, you must configure the server runtime environment.

Environment Variables

Set these required variables defined in infra/config/config.go:

  • GROK2API_HTTP_PORT: Port for the HTTP server (e.g., 8080)
  • GROK2API_JWT_SIGNING_KEY: Secret key for JWT signing and validation
  • GROK2API_REDIS_URL: Optional Redis connection URL; defaults to in-memory storage if omitted

Docker Deployment

Run Grok2API as a standalone service using the provided Dockerfile in the repository root:

docker build -t grok2api .
docker run -e GROK2API_HTTP_PORT=8080 -e GROK2API_JWT_SIGNING_KEY=your-secret -p 8080:8080 grok2api

Alternatively, run directly from source:

go run ./backend/cmd/grok2api

Authentication Flow

Grok2API uses JWT-based authentication implemented in infra/security/token.go.

  1. Obtain a token by calling POST /v1/account/login with valid credentials
  2. Include the token in subsequent requests as Authorization: Bearer <token>

The transport layer validates tokens before routing to handlers in transport/http/inference/handler.go and other endpoints.

Core Integration Patterns

You can integrate Grok2API using three primary patterns depending on your existing architecture.

HTTP REST API Consumption

The most flexible approach uses standard HTTP clients. The server exposes routes defined in transport/http/server.go, including:

  • /v1/inference - Text generation and chat completion
  • /v1/models - Model management (GET, POST, DELETE)
  • /v1/media/upload and /v1/media/{id} - Media file handling
  • /v1/dashboard/* and /v1/audit/* - Usage statistics and logging

Embedding as a Go Module

For Go applications, import the service layer directly from application/model/service.go and domain entities from domain/model/model.go. This bypasses HTTP transport and calls business logic directly, though you must manually wire dependencies like the runtime store (infra/runtime/memory/store.go or infra/runtime/redis/store.go) and provider (infra/provider/web/chat.go).

Provider Integration

The infra/provider/web/ package handles outbound calls to external LLM services. When embedding Grok2API, you can inject custom providers by implementing the interfaces defined in the domain layer, allowing you to route inference requests to private models or specialized hardware.

Working with Key Endpoints

Inference Requests

Send prompts to POST /v1/inference. The handler in transport/http/inference/handler.go forwards requests to the provider implementation in infra/provider/web/chat.go, which communicates with external LLM APIs.

Request structure:

{
  "prompt": "Explain quantum computing",
  "model_id": "gpt-4o-mini",
  "max_tokens": 200
}

Model Management

CRUD operations for models route through transport/http/model/handler.go and use the application service in application/model/service.go. Domain entities are defined in domain/model/model.go.

Media Handling

Upload files via POST /v1/media/upload (handled in transport/http/media/handler.go) and reference them in inference requests using the media ID. Assets are modeled in domain/media/asset.go.

Integration Code Examples

Go HTTP Client Example

Call the inference endpoint from an existing Go application:

package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"net/http"
)

type InferenceRequest struct {
	Prompt   string `json:"prompt"`
	ModelID  string `json:"model_id"`
	MaxTokens int   `json:"max_tokens,omitempty"`
}

type InferenceResponse struct {
	Answer string `json:"answer"`
}

func main() {
	// Obtain JWT via your actual login flow
	jwtToken := "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9..."

	reqBody := InferenceRequest{
		Prompt:   "Explain the difference between REST and GraphQL.",
		ModelID:  "gpt-4o-mini",
		MaxTokens: 200,
	}
	data, _ := json.Marshal(reqBody)

	req, _ := http.NewRequest("POST", "http://localhost:8080/v1/inference", bytes.NewBuffer(data))
	req.Header.Set("Content-Type", "application/json")
	req.Header.Set("Authorization", "Bearer "+jwtToken)

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()

	var out InferenceResponse
	if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
		panic(err)
	}
	fmt.Println("Answer:", out.Answer)
}

Bash Media Upload Workflow

Upload media and reference it in inference using curl:


# Upload file and extract media ID

MEDIA_ID=$(curl -s -X POST -H "Authorization: Bearer $JWT" \
  -F "file=@/path/to/image.png" http://localhost:8080/v1/media/upload | jq -r .id)

# Submit inference with media reference

curl -X POST -H "Authorization: Bearer $JWT" -H "Content-Type: application/json" \
  -d '{"prompt":"Describe the image", "model_id":"gpt-4o-mini", "media_ids":["'"$MEDIA_ID"'"]}' \
  http://localhost:8080/v1/inference

Summary

  • Grok2API follows clean architecture with clear separation between transport, application, domain, and infrastructure layers in backend/internal/.
  • Configuration requires setting GROK2API_HTTP_PORT, GROK2API_JWT_SIGNING_KEY, and optionally GROK2API_REDIS_URL in the environment.
  • Authentication uses JWT tokens obtained from POST /v1/account/login and validated via infra/security/token.go.
  • Integration options include HTTP REST calls, direct Go module imports from application/ and domain/ packages, or Docker containerization.
  • Key endpoints cover inference (/v1/inference), model management (/v1/models), and media handling (/v1/media/*), with handlers orchestrating calls to the web provider in infra/provider/web/chat.go.

Frequently Asked Questions

How do I configure Grok2API to use Redis instead of in-memory storage?

Set the GROK2API_REDIS_URL environment variable to your Redis connection string (e.g., redis://localhost:6379). If omitted, the server defaults to the in-memory runtime store implemented in infra/runtime/memory/store.go. The Redis implementation in infra/runtime/redis/store.go handles connection pooling and key management automatically.

Can I call Grok2API endpoints from a Python or JavaScript application?

Yes. Grok2API exposes standard RESTful HTTP endpoints that accept JSON payloads and return JSON responses. Use any HTTP client library (such as Python's requests or JavaScript's fetch) to send authenticated requests to http://localhost:8080/v1/inference or other routes. Include the JWT token in the Authorization: Bearer <token> header as implemented in the transport layer handlers.

What is the difference between the web provider and the application service layer?

The application service layer in application/model/service.go contains business logic and orchestration rules, while the web provider in infra/provider/web/chat.go handles the actual outbound HTTP calls to external LLM APIs. The application layer calls the provider through domain interfaces, maintaining clean architecture boundaries and allowing you to swap providers without changing business logic.

How do I extend Grok2API to support custom authentication methods?

Modify the security infrastructure in infra/security/token.go to implement your custom validation logic, or add new middleware in the transport layer (transport/http/server.go). The clean architecture allows you to inject custom security implementations into the handler registration flow without affecting domain or application logic. Ensure your custom implementation satisfies the token interfaces defined in the domain layer.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →