How to Configure Client Keys with Model Allowlists and Rate Limits in Grok2API

Grok2API enables fine-grained access control by configuring client keys with specific model allowlists, requests-per-minute (RPM) caps, and concurrency limits through the ClientKeyRepository interface, with enforcement handled by sharded in-memory limiters in the authentication middleware.

Grok2API is a Go-based gateway that unifies Grok Build, Grok Web, and Grok Console services behind an OpenAI-compatible API endpoint. When you configure client keys with model allowlists and rate limits, you create isolated tenant boundaries that prevent resource abuse and ensure fair access across teams. This guide walks through the domain model, repository layer, and runtime enforcement mechanisms implemented in the chenyme/grok2api repository.

Client Key Architecture and Domain Model

The platform organizes functionality into three logical domains: Access (API clients and admin UI), Core (management and routing), and Provider (upstream adapters). Client keys are persisted and managed through the ClientKeyRepository interface defined in backend/internal/repository/client_key.go:

type ClientKeyRepository interface {
    List(ctx context.Context, query ClientKeyListQuery) ([]clientkey.Key, int64, error)
    Create(ctx context.Context, value clientkey.Key) (clientkey.Key, error)
    Get(ctx context.Context, id uint64) (clientkey.Key, error)
    GetByPrefix(ctx context.Context, prefix string) (clientkey.Key, error)
    Update(ctx context.Context, value clientkey.Key) (clientkey.Key, error)
    UpdateManyEnabled(ctx context.Context, ids []uint64, enabled bool) (int64, error)
    Delete(ctx context.Context, id uint64) error
    DeleteMany(ctx context.Context, ids []uint64) (int64, error)
    Touch(ctx context.Context, id uint64) error
}

The domain model in backend/internal/domain/clientkey/key.go defines the key structure that supports granular access control. Each key contains a Prefix (the first 8-12 characters used for fast lookups), an Allowlist of permitted model identifiers, a RateLimitRPM ceiling, and a ConcurrencyLimit for simultaneous requests.

Defining Allowlists and Rate Limits

When you configure client keys with model allowlists and rate limits, you populate four critical fields in the clientkey.Key struct:

  • Prefix – The public identifier used to resolve the full key configuration without exposing the secret
  • Allowlist – A slice of model strings (e.g., grok-4.0, grok-4-code) that the key is authorized to access
  • RateLimitRPM – The maximum number of requests allowed per minute
  • ConcurrencyLimit – The maximum number of simultaneous in-flight requests

These values are persisted in SQLite or PostgreSQL via GORM and loaded into memory at runtime for O(1) enforcement checks.

Runtime Enforcement Mechanisms

The authentication middleware in backend/internal/transport/http/middleware/auth.go orchestrates enforcement by extracting the Authorization: Bearer <key> header and resolving the configuration via clientkeyapp.GetByPrefix.

Rate Limiting (RPM)

The RateLimiter in backend/internal/infra/runtime/memory/store.go implements a fixed-minute-window algorithm partitioned across 64 shards to eliminate lock contention:

type RateLimiter struct {
    shards [shardCount]rateShard
}

When a request arrives, the Allow method checks the current minute window count against the key’s RateLimitRPM. If the limit is exceeded, the middleware returns 429 Too Many Requests with a retry_after value indicating seconds until the next window.

Concurrency Limiting

The ConcurrencyLimiter tracks active requests per key in the same memory store. It atomically increments a counter when a request enters and decrements on exit. If the count reaches the key’s ConcurrencyLimit, the gateway returns 503 Service Unavailable with the error type concurrencyLimited.

Model Allowlist Validation

After resolving the key and passing limit checks, the middleware validates the request payload’s model field against the key’s Allowlist slice. If the model is not permitted, the gateway responds with 403 Forbidden and the error modelNotAllowed.

Configuration and Implementation Examples

YAML Configuration

Define keys statically in config.example.yaml under the clientKeys section:

clientKeys:
  - prefix: "abcd1234"
    enabled: true
    allowlist:
      - "grok-4.0"
      - "grok-4-code"
    rateLimitRPM: 120
    concurrencyLimit: 5
    description: "Team Alpha key"

Programmatic Key Creation

Use the repository interface to create keys dynamically:

import (
    "context"
    "github.com/chenyme/grok2api/backend/internal/domain/clientkey"
    "github.com/chenyme/grok2api/backend/internal/repository"
)

func CreateRestrictedKey(repo repository.ClientKeyRepository) (*clientkey.Key, error) {
    key := clientkey.Key{
        Prefix:           "teamalpha",
        Enabled:          true,
        Allowlist:        []string{"grok-4.0", "grok-4-code"},
        RateLimitRPM:    120,
        ConcurrencyLimit: 5,
        Description:      "Production workload key",
    }
    return repo.Create(context.Background(), key)
}

API Usage and Error Handling

Send requests with the Bearer token:

curl -X POST http://localhost:8000/v1/chat/completions \
     -H "Authorization: Bearer abcd1234xyz..." \
     -H "Content-Type: application/json" \
     -d '{"model": "grok-4.0", "messages": [{"role":"user","content":"Hello"}]}'

If the model is not in the allowlist:

{
  "error": {
    "type": "access_error",
    "message": "model_not_allowed"
  }
}

When RPM limits are exceeded:

{
  "error": {
    "type": "rate_limit_error",
    "message": "rate_limit_exceeded",
    "retry_after": 30
  }
}

Summary

  • Client keys in Grok2API are managed through the ClientKeyRepository interface in backend/internal/repository/client_key.go, supporting CRUD operations and batch updates.
  • Model allowlists are enforced by validating the request model against the Allowlist field in backend/internal/transport/http/middleware/auth.go.
  • Rate limits (RPM) use a sharded in-memory limiter in backend/internal/infra/runtime/memory/store.go to track per-minute request counts and return 429 errors when exceeded.
  • Concurrency limits are enforced by the same store component, returning 503 errors when simultaneous request caps are reached.
  • Configuration can be defined statically in YAML or dynamically via the repository API, with the React admin UI providing a visual management interface.

Frequently Asked Questions

What happens when a client exceeds the RPM limit?

The RateLimiter in backend/internal/infra/runtime/memory/store.go tracks requests in fixed one-minute windows. When the count exceeds the key’s RateLimitRPM, the authentication middleware returns 429 Too Many Requests with a retry_after field indicating how many seconds remain until the next minute window resets. Clients should implement exponential backoff based on this value.

Can I update rate limits without restarting the service?

Yes. The ClientKeyRepository.Update method in backend/internal/repository/client_key.go persists changes to the database, and the Touch method can trigger runtime cache refreshes. The admin UI (frontend/src/...) provides a web interface for modifying RateLimitRPM and ConcurrencyLimit values dynamically without service restarts.

How does the concurrency limiter prevent race conditions?

The ConcurrencyLimiter uses atomic operations across 64 shards to track active requests per key. When a request enters the handler in backend/internal/transport/http/inference/handler.go, the limiter atomically increments the counter; on completion (success or error), it decrements. This lock-sharded design ensures thread-safe enforcement without blocking the request path under high load.

What is the performance impact of the allowlist check?

The allowlist validation performs an O(n) slice lookup against the key’s Allowlist field, where n is typically small (under 10 models per key). This check occurs after the O(1) prefix lookup in GetByPrefix and before routing to the provider adapter, adding negligible latency to the request lifecycle.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →