# How to Configure Client Keys with Model Allowlists and Rate Limits in Grok2API

> Learn how to configure client keys with model allowlists and rate limits in Grok2API. Control access with RPM and concurrency limits for secure API management.

- Repository: [Chenyme/grok2api](https://github.com/chenyme/grok2api)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Grok2API enables fine-grained access control by configuring client keys with specific model allowlists, requests-per-minute (RPM) caps, and concurrency limits through the `ClientKeyRepository` interface, with enforcement handled by sharded in-memory limiters in the authentication middleware.**

Grok2API is a Go-based gateway that unifies Grok Build, Grok Web, and Grok Console services behind an OpenAI-compatible API endpoint. When you configure client keys with model allowlists and rate limits, you create isolated tenant boundaries that prevent resource abuse and ensure fair access across teams. This guide walks through the domain model, repository layer, and runtime enforcement mechanisms implemented in the `chenyme/grok2api` repository.

## Client Key Architecture and Domain Model

The platform organizes functionality into three logical domains: **Access** (API clients and admin UI), **Core** (management and routing), and **Provider** (upstream adapters). Client keys are persisted and managed through the `ClientKeyRepository` interface defined in [`backend/internal/repository/client_key.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/repository/client_key.go):

```go
type ClientKeyRepository interface {
    List(ctx context.Context, query ClientKeyListQuery) ([]clientkey.Key, int64, error)
    Create(ctx context.Context, value clientkey.Key) (clientkey.Key, error)
    Get(ctx context.Context, id uint64) (clientkey.Key, error)
    GetByPrefix(ctx context.Context, prefix string) (clientkey.Key, error)
    Update(ctx context.Context, value clientkey.Key) (clientkey.Key, error)
    UpdateManyEnabled(ctx context.Context, ids []uint64, enabled bool) (int64, error)
    Delete(ctx context.Context, id uint64) error
    DeleteMany(ctx context.Context, ids []uint64) (int64, error)
    Touch(ctx context.Context, id uint64) error
}

```

The domain model in [`backend/internal/domain/clientkey/key.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/domain/clientkey/key.go) defines the key structure that supports granular access control. Each key contains a **Prefix** (the first 8-12 characters used for fast lookups), an **Allowlist** of permitted model identifiers, a **RateLimitRPM** ceiling, and a **ConcurrencyLimit** for simultaneous requests.

## Defining Allowlists and Rate Limits

When you configure client keys with model allowlists and rate limits, you populate four critical fields in the `clientkey.Key` struct:

- **Prefix** – The public identifier used to resolve the full key configuration without exposing the secret
- **Allowlist** – A slice of model strings (e.g., `grok-4.0`, `grok-4-code`) that the key is authorized to access
- **RateLimitRPM** – The maximum number of requests allowed per minute
- **ConcurrencyLimit** – The maximum number of simultaneous in-flight requests

These values are persisted in SQLite or PostgreSQL via GORM and loaded into memory at runtime for O(1) enforcement checks.

## Runtime Enforcement Mechanisms

The authentication middleware in [`backend/internal/transport/http/middleware/auth.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/middleware/auth.go) orchestrates enforcement by extracting the `Authorization: Bearer <key>` header and resolving the configuration via `clientkeyapp.GetByPrefix`.

### Rate Limiting (RPM)

The **RateLimiter** in [`backend/internal/infra/runtime/memory/store.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/runtime/memory/store.go) implements a fixed-minute-window algorithm partitioned across 64 shards to eliminate lock contention:

```go
type RateLimiter struct {
    shards [shardCount]rateShard
}

```

When a request arrives, the `Allow` method checks the current minute window count against the key’s `RateLimitRPM`. If the limit is exceeded, the middleware returns **429 Too Many Requests** with a `retry_after` value indicating seconds until the next window.

### Concurrency Limiting

The **ConcurrencyLimiter** tracks active requests per key in the same memory store. It atomically increments a counter when a request enters and decrements on exit. If the count reaches the key’s `ConcurrencyLimit`, the gateway returns **503 Service Unavailable** with the error type `concurrencyLimited`.

### Model Allowlist Validation

After resolving the key and passing limit checks, the middleware validates the request payload’s `model` field against the key’s `Allowlist` slice. If the model is not permitted, the gateway responds with **403 Forbidden** and the error `modelNotAllowed`.

## Configuration and Implementation Examples

### YAML Configuration

Define keys statically in [`config.example.yaml`](https://github.com/chenyme/grok2api/blob/main/config.example.yaml) under the `clientKeys` section:

```yaml
clientKeys:
  - prefix: "abcd1234"
    enabled: true
    allowlist:
      - "grok-4.0"
      - "grok-4-code"
    rateLimitRPM: 120
    concurrencyLimit: 5
    description: "Team Alpha key"

```

### Programmatic Key Creation

Use the repository interface to create keys dynamically:

```go
import (
    "context"
    "github.com/chenyme/grok2api/backend/internal/domain/clientkey"
    "github.com/chenyme/grok2api/backend/internal/repository"
)

func CreateRestrictedKey(repo repository.ClientKeyRepository) (*clientkey.Key, error) {
    key := clientkey.Key{
        Prefix:           "teamalpha",
        Enabled:          true,
        Allowlist:        []string{"grok-4.0", "grok-4-code"},
        RateLimitRPM:    120,
        ConcurrencyLimit: 5,
        Description:      "Production workload key",
    }
    return repo.Create(context.Background(), key)
}

```

### API Usage and Error Handling

Send requests with the Bearer token:

```bash
curl -X POST http://localhost:8000/v1/chat/completions \
     -H "Authorization: Bearer abcd1234xyz..." \
     -H "Content-Type: application/json" \
     -d '{"model": "grok-4.0", "messages": [{"role":"user","content":"Hello"}]}'

```

If the model is not in the allowlist:

```json
{
  "error": {
    "type": "access_error",
    "message": "model_not_allowed"
  }
}

```

When RPM limits are exceeded:

```json
{
  "error": {
    "type": "rate_limit_error",
    "message": "rate_limit_exceeded",
    "retry_after": 30
  }
}

```

## Summary

- **Client keys** in Grok2API are managed through the `ClientKeyRepository` interface in [`backend/internal/repository/client_key.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/repository/client_key.go), supporting CRUD operations and batch updates.
- **Model allowlists** are enforced by validating the request model against the `Allowlist` field in [`backend/internal/transport/http/middleware/auth.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/middleware/auth.go).
- **Rate limits (RPM)** use a sharded in-memory limiter in [`backend/internal/infra/runtime/memory/store.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/runtime/memory/store.go) to track per-minute request counts and return **429** errors when exceeded.
- **Concurrency limits** are enforced by the same store component, returning **503** errors when simultaneous request caps are reached.
- Configuration can be defined statically in YAML or dynamically via the repository API, with the React admin UI providing a visual management interface.

## Frequently Asked Questions

### What happens when a client exceeds the RPM limit?

The **RateLimiter** in [`backend/internal/infra/runtime/memory/store.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/infra/runtime/memory/store.go) tracks requests in fixed one-minute windows. When the count exceeds the key’s `RateLimitRPM`, the authentication middleware returns **429 Too Many Requests** with a `retry_after` field indicating how many seconds remain until the next minute window resets. Clients should implement exponential backoff based on this value.

### Can I update rate limits without restarting the service?

Yes. The `ClientKeyRepository.Update` method in [`backend/internal/repository/client_key.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/repository/client_key.go) persists changes to the database, and the `Touch` method can trigger runtime cache refreshes. The admin UI (`frontend/src/...`) provides a web interface for modifying `RateLimitRPM` and `ConcurrencyLimit` values dynamically without service restarts.

### How does the concurrency limiter prevent race conditions?

The **ConcurrencyLimiter** uses atomic operations across 64 shards to track active requests per key. When a request enters the handler in [`backend/internal/transport/http/inference/handler.go`](https://github.com/chenyme/grok2api/blob/main/backend/internal/transport/http/inference/handler.go), the limiter atomically increments the counter; on completion (success or error), it decrements. This lock-sharded design ensures thread-safe enforcement without blocking the request path under high load.

### What is the performance impact of the allowlist check?

The allowlist validation performs an O(n) slice lookup against the key’s `Allowlist` field, where n is typically small (under 10 models per key). This check occurs after the O(1) prefix lookup in `GetByPrefix` and before routing to the provider adapter, adding negligible latency to the request lifecycle.