# How to Integrate Grok2API with an Existing Application: Complete Technical Guide

> Learn to integrate Grok2API with your application. This guide covers server setup, JWT authentication, and consuming RESTful endpoints for inference and media handling. Get started today.

- Repository: [Chenyme/grok2api](https://github.com/chenyme/grok2api)
- Tags: how-to-guide
- Published: 2026-07-16

---

**To integrate Grok2API with an existing application, run the Go-based server using environment variables for configuration, obtain JWT tokens via the authentication endpoint, and consume RESTful endpoints for inference, model management, and media handling.**

Grok2API is a production-ready Go server that exposes a RESTful HTTP interface for AI inference and model orchestration. According to the chenyme/grok2api source code, the system follows hexagonal (clean) architecture principles organized into distinct transport, application, domain, and infrastructure layers. This guide demonstrates how to integrate Grok2API with an existing application using HTTP clients, Go module imports, or Docker containerization.

## Understanding the Grok2API Architecture

Grok2API implements a **hexagonal architecture** that separates concerns into distinct layers, making integration straightforward regardless of your tech stack.

### Architecture Layers

The codebase in `backend/internal/` follows these clear boundaries:

- **Transport Layer** (`transport/http/*`): Handles HTTP routing, authentication, and request/response formatting. Key sub-packages include `inference`, `model`, `media`, `dashboard`, and `audit`.
- **Application Layer** (`application/*`): Contains business logic and service implementations such as `settings`, `quotarecovery`, and `model` services.
- **Domain Layer** (`domain/*`): Defines pure business entities including `model`, `media`, `inference`, and `clientkey`.
- **Infrastructure Layer** (`infra/*`): Provides concrete implementations for external concerns including `security` (JWT), `egress` (outbound HTTP), `runtime` (storage), and `provider` (web/console adapters).

The server bootstrap in [`backend/cmd/grok2api/main.go`](https://github.com/chenyme/grok2api/blob/main/backend/cmd/grok2api/main.go) wires these components together, creating a dependency injection container that initializes the HTTP server.

## Configuration and Deployment Options

Before integrating, you must configure the server runtime environment.

### Environment Variables

Set these required variables defined in [`infra/config/config.go`](https://github.com/chenyme/grok2api/blob/main/infra/config/config.go):

- `GROK2API_HTTP_PORT`: Port for the HTTP server (e.g., `8080`)
- `GROK2API_JWT_SIGNING_KEY`: Secret key for JWT signing and validation
- `GROK2API_REDIS_URL`: Optional Redis connection URL; defaults to in-memory storage if omitted

### Docker Deployment

Run Grok2API as a standalone service using the provided Dockerfile in the repository root:

```bash
docker build -t grok2api .
docker run -e GROK2API_HTTP_PORT=8080 -e GROK2API_JWT_SIGNING_KEY=your-secret -p 8080:8080 grok2api

```

Alternatively, run directly from source:

```bash
go run ./backend/cmd/grok2api

```

## Authentication Flow

Grok2API uses **JWT-based authentication** implemented in [`infra/security/token.go`](https://github.com/chenyme/grok2api/blob/main/infra/security/token.go).

1. **Obtain a token** by calling `POST /v1/account/login` with valid credentials
2. **Include the token** in subsequent requests as `Authorization: Bearer <token>`

The transport layer validates tokens before routing to handlers in [`transport/http/inference/handler.go`](https://github.com/chenyme/grok2api/blob/main/transport/http/inference/handler.go) and other endpoints.

## Core Integration Patterns

You can integrate Grok2API using three primary patterns depending on your existing architecture.

### HTTP REST API Consumption

The most flexible approach uses standard HTTP clients. The server exposes routes defined in [`transport/http/server.go`](https://github.com/chenyme/grok2api/blob/main/transport/http/server.go), including:

- `/v1/inference` - Text generation and chat completion
- `/v1/models` - Model management (GET, POST, DELETE)
- `/v1/media/upload` and `/v1/media/{id}` - Media file handling
- `/v1/dashboard/*` and `/v1/audit/*` - Usage statistics and logging

### Embedding as a Go Module

For Go applications, import the service layer directly from [`application/model/service.go`](https://github.com/chenyme/grok2api/blob/main/application/model/service.go) and domain entities from [`domain/model/model.go`](https://github.com/chenyme/grok2api/blob/main/domain/model/model.go). This bypasses HTTP transport and calls business logic directly, though you must manually wire dependencies like the runtime store ([`infra/runtime/memory/store.go`](https://github.com/chenyme/grok2api/blob/main/infra/runtime/memory/store.go) or [`infra/runtime/redis/store.go`](https://github.com/chenyme/grok2api/blob/main/infra/runtime/redis/store.go)) and provider ([`infra/provider/web/chat.go`](https://github.com/chenyme/grok2api/blob/main/infra/provider/web/chat.go)).

### Provider Integration

The `infra/provider/web/` package handles outbound calls to external LLM services. When embedding Grok2API, you can inject custom providers by implementing the interfaces defined in the domain layer, allowing you to route inference requests to private models or specialized hardware.

## Working with Key Endpoints

### Inference Requests

Send prompts to `POST /v1/inference`. The handler in [`transport/http/inference/handler.go`](https://github.com/chenyme/grok2api/blob/main/transport/http/inference/handler.go) forwards requests to the provider implementation in [`infra/provider/web/chat.go`](https://github.com/chenyme/grok2api/blob/main/infra/provider/web/chat.go), which communicates with external LLM APIs.

**Request structure:**

```json
{
  "prompt": "Explain quantum computing",
  "model_id": "gpt-4o-mini",
  "max_tokens": 200
}

```

### Model Management

CRUD operations for models route through [`transport/http/model/handler.go`](https://github.com/chenyme/grok2api/blob/main/transport/http/model/handler.go) and use the application service in [`application/model/service.go`](https://github.com/chenyme/grok2api/blob/main/application/model/service.go). Domain entities are defined in [`domain/model/model.go`](https://github.com/chenyme/grok2api/blob/main/domain/model/model.go).

### Media Handling

Upload files via `POST /v1/media/upload` (handled in [`transport/http/media/handler.go`](https://github.com/chenyme/grok2api/blob/main/transport/http/media/handler.go)) and reference them in inference requests using the media ID. Assets are modeled in [`domain/media/asset.go`](https://github.com/chenyme/grok2api/blob/main/domain/media/asset.go).

## Integration Code Examples

### Go HTTP Client Example

Call the inference endpoint from an existing Go application:

```go
package main

import (
	"bytes"
	"encoding/json"
	"fmt"
	"net/http"
)

type InferenceRequest struct {
	Prompt   string `json:"prompt"`
	ModelID  string `json:"model_id"`
	MaxTokens int   `json:"max_tokens,omitempty"`
}

type InferenceResponse struct {
	Answer string `json:"answer"`
}

func main() {
	// Obtain JWT via your actual login flow
	jwtToken := "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9..."

	reqBody := InferenceRequest{
		Prompt:   "Explain the difference between REST and GraphQL.",
		ModelID:  "gpt-4o-mini",
		MaxTokens: 200,
	}
	data, _ := json.Marshal(reqBody)

	req, _ := http.NewRequest("POST", "http://localhost:8080/v1/inference", bytes.NewBuffer(data))
	req.Header.Set("Content-Type", "application/json")
	req.Header.Set("Authorization", "Bearer "+jwtToken)

	resp, err := http.DefaultClient.Do(req)
	if err != nil {
		panic(err)
	}
	defer resp.Body.Close()

	var out InferenceResponse
	if err := json.NewDecoder(resp.Body).Decode(&out); err != nil {
		panic(err)
	}
	fmt.Println("Answer:", out.Answer)
}

```

### Bash Media Upload Workflow

Upload media and reference it in inference using `curl`:

```bash

# Upload file and extract media ID

MEDIA_ID=$(curl -s -X POST -H "Authorization: Bearer $JWT" \
  -F "file=@/path/to/image.png" http://localhost:8080/v1/media/upload | jq -r .id)

# Submit inference with media reference

curl -X POST -H "Authorization: Bearer $JWT" -H "Content-Type: application/json" \
  -d '{"prompt":"Describe the image", "model_id":"gpt-4o-mini", "media_ids":["'"$MEDIA_ID"'"]}' \
  http://localhost:8080/v1/inference

```

## Summary

- **Grok2API** follows clean architecture with clear separation between transport, application, domain, and infrastructure layers in `backend/internal/`.
- **Configuration** requires setting `GROK2API_HTTP_PORT`, `GROK2API_JWT_SIGNING_KEY`, and optionally `GROK2API_REDIS_URL` in the environment.
- **Authentication** uses JWT tokens obtained from `POST /v1/account/login` and validated via [`infra/security/token.go`](https://github.com/chenyme/grok2api/blob/main/infra/security/token.go).
- **Integration options** include HTTP REST calls, direct Go module imports from `application/` and `domain/` packages, or Docker containerization.
- **Key endpoints** cover inference (`/v1/inference`), model management (`/v1/models`), and media handling (`/v1/media/*`), with handlers orchestrating calls to the web provider in [`infra/provider/web/chat.go`](https://github.com/chenyme/grok2api/blob/main/infra/provider/web/chat.go).

## Frequently Asked Questions

### How do I configure Grok2API to use Redis instead of in-memory storage?

Set the `GROK2API_REDIS_URL` environment variable to your Redis connection string (e.g., `redis://localhost:6379`). If omitted, the server defaults to the in-memory runtime store implemented in [`infra/runtime/memory/store.go`](https://github.com/chenyme/grok2api/blob/main/infra/runtime/memory/store.go). The Redis implementation in [`infra/runtime/redis/store.go`](https://github.com/chenyme/grok2api/blob/main/infra/runtime/redis/store.go) handles connection pooling and key management automatically.

### Can I call Grok2API endpoints from a Python or JavaScript application?

Yes. Grok2API exposes standard RESTful HTTP endpoints that accept JSON payloads and return JSON responses. Use any HTTP client library (such as Python's `requests` or JavaScript's `fetch`) to send authenticated requests to `http://localhost:8080/v1/inference` or other routes. Include the JWT token in the `Authorization: Bearer <token>` header as implemented in the transport layer handlers.

### What is the difference between the web provider and the application service layer?

The **application service** layer in [`application/model/service.go`](https://github.com/chenyme/grok2api/blob/main/application/model/service.go) contains business logic and orchestration rules, while the **web provider** in [`infra/provider/web/chat.go`](https://github.com/chenyme/grok2api/blob/main/infra/provider/web/chat.go) handles the actual outbound HTTP calls to external LLM APIs. The application layer calls the provider through domain interfaces, maintaining clean architecture boundaries and allowing you to swap providers without changing business logic.

### How do I extend Grok2API to support custom authentication methods?

Modify the security infrastructure in [`infra/security/token.go`](https://github.com/chenyme/grok2api/blob/main/infra/security/token.go) to implement your custom validation logic, or add new middleware in the transport layer ([`transport/http/server.go`](https://github.com/chenyme/grok2api/blob/main/transport/http/server.go)). The clean architecture allows you to inject custom security implementations into the handler registration flow without affecting domain or application logic. Ensure your custom implementation satisfies the token interfaces defined in the domain layer.