# How to Scale AI-Infra-Guard in a Production Environment: 7-Step Deployment Guide

> Scale AI-Infra-Guard in production with this 7-step guide. Learn to containerize, deploy replicas, and use external databases for robust performance.

- Repository: [Tencent/AI-Infra-Guard](https://github.com/tencent/AI-Infra-Guard)
- Tags: how-to-guide
- Published: 2026-08-23

---

**Deploy AI-Infra-Guard at scale by containerizing the Go-based server and lightweight agents, running multiple replicas behind a load balancer, and using external PostgreSQL or Redis for shared state persistence.**

AI-Infra-Guard (AIG) is an open-source AI infrastructure scanner built by Tencent that combines a modular Go server with lightweight agent workers to audit machine learning environments. The repository (`Tencent/AI-Infra-Guard) provides production-ready Docker images, WebSocket-based task distribution, and built-in Prometheus metrics, making it straightforward to scale from a single-node deployment to a high-availability cluster. This guide explains how to configure the core components—defined in [`cmd/cli/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go) and [`common/websocket/task_manager.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/task_manager.go)—for enterprise-grade workloads.

## Architecture Overview

Before scaling, understand how the components interact in a distributed setup. The platform consists of four primary layers:

- **Server Layer** ([`cmd/cli/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go)): The Go HTTP/WebSocket API that accepts scan requests, schedules tasks, and aggregates results.
- **Agent Workers** (`Dockerfile_Agent`): Lightweight containers that register via WebSocket ([`common/websocket/websocket.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/websocket.go)) and execute I/O-bound scans.
- **State Store**: External PostgreSQL or Redis database configured via `AIG_DB_URL` to persist task states across server restarts.
- **Load Balancer**: Distributes incoming HTTPS and WebSocket traffic across multiple server replicas.

The [`common/websocket/task_manager.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/task_manager.go) file implements the core queue logic, allowing agents to pick up pending jobs concurrently without blocking the main server process.

## Step-by-Step Production Scaling

### 1. Containerize with Official Images

Start by using the pre-built images or building from the Dockerfiles in the repository. The `Dockerfile` creates the main AIG server containing the Go binary and web UI, while `Dockerfile_Agent` produces the minimal agent image used by MCP/Skill scanners.

Guarantee reproducible environments by pinning to specific image tags rather than `latest`. This isolation prevents dependency conflicts when running mixed versions across your cluster.

### 2. Implement Horizontal Scaling of Server Replicas

Deploy multiple instances of the server behind a load balancer such as Nginx, HAProxy, or a Kubernetes Service. Each replica in [`cmd/cli/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go) operates independently, accepting scan requests via the HTTP API while the WebSocket handler manages long-lived connections to agents.

Configure your load balancer to support **sticky sessions** for WebSocket connections or use a stateless round-robin for the REST API endpoints. This distribution prevents any single node from becoming a bottleneck during peak scanning periods.

### 3. Configure Shared State Persistence

Prevent data loss during pod restarts by externalizing the database. The server supports PostgreSQL, Redis, or mounted volumes via configuration flags defined in [`cmd/cli/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go).

Set the `AIG_DB_URL` environment variable to point to your external store:

```bash
AIG_DB_URL=postgres://aig:password@db:5432/aig?sslmode=disable

```

This ensures that task states and scan results survive server crashes and allow newly spawned workers to resume pending jobs immediately.

### 4. Deploy Dedicated Agent Workers

Scale scan throughput by running multiple `agent` containers built from `Dockerfile_Agent`. These workers connect to the main server via WebSocket ([`common/websocket/websocket.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/websocket.go)), register themselves automatically, and pull tasks from the distributed queue managed by [`common/websocket/task_manager.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/task_manager.go).

Since scans are I/O-bound rather than CPU-bound, you can run many agent replicas on modest hardware. Each agent operates independently, so adding containers linearly increases throughput without overloading the server process.

### 5. Integrate the API-Checker Service

For model fingerprinting and relay checking capabilities, deploy the separate `api-checker` service located in `services/api_checker`. Configure the `AIG_API_CHECKER_PYTHON` environment variable as documented in the repository README.

Keeping this component isolated allows you to scale Python-based security checks independently from the Go server, preventing resource contention during intensive model analysis.

### 6. Enable Monitoring and Autoscaling

The Go server exposes Prometheus metrics at the `/metrics` endpoint when started with the `--metrics` flag or `AIG_ENABLE_METRICS=true` environment variable.

Configure your monitoring stack to scrape:

```bash
AIG_ENABLE_METRICS=true ./ai-infra-guard webserver --server 0.0.0.0:8088

```

Set up autoscaling rules based on CPU utilization, memory consumption, and queue depth metrics. This guarantees SLA compliance and automatic horizontal scaling during traffic spikes.

### 7. Secure the Deployment

AI-Infra-Guard **has no built-in authentication**. Always deploy inside a trusted network boundary (VPC or VPN) and restrict inbound ports to internal services only. Terminate TLS at the edge load balancer to encrypt WebSocket and HTTP traffic.

Never expose the scan APIs directly to the public internet, as this could allow unauthorized access to your infrastructure auditing capabilities.

## Production Deployment Examples

### Docker Compose with Replicas

The [`docker-compose.images.yml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/docker-compose.images.yml) file provides a template for production deployment. Scale the server and agents independently using the `deploy.replicas` directive:

```yaml
version: '3.8'
services:
  aig-server:
    image: zhuquelab/aig-server:latest
    restart: always
    environment:
      - AIG_DB_URL=postgres://aig:password@db:5432/aig?sslmode=disable
    ports:
      - "8088:8088"
    depends_on:
      - db
    deploy:
      mode: replicated
      replicas: 4
      resources:
        limits:
          cpus: "2"
          memory: 4G

  aig-agent:
    image: zhuquelab/aig-agent:latest
    restart: always
    environment:
      - AIG_SERVER_URL=http://aig-server:8088
    deploy:
      mode: replicated
      replicas: 8
      resources:
        limits:
          cpus: "1"
          memory: 2G

  db:
    image: postgres:15
    restart: always
    environment:
      POSTGRES_USER: aig
      POSTGRES_PASSWORD: password
      POSTGRES_DB: aig
    volumes:
      - db-data:/var/lib/postgresql/data

volumes:
  db-data:

```

### Kubernetes Deployment

For orchestrated environments, separate the server and agent into distinct Deployments. The server handles API requests while agents scale based on scan backlog:

```yaml
apiVersion: apps/v1
kind: Deployment
metadata:
  name: aig-server
spec:
  replicas: 3
  selector:
    matchLabels:
      app: aig-server
  template:
    metadata:
      labels:
        app: aig-server
    spec:
      containers:
        - name: server
          image: zhuquelab/aig-server:latest
          ports:
            - containerPort: 8088
          env:
            - name: AIG_DB_URL
              value: postgres://aig:password@postgres:5432/aig?sslmode=disable
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: aig-agent
spec:
  replicas: 6
  selector:
    matchLabels:
      app: aig-agent
  template:
    metadata:
      labels:
        app: aig-agent
    spec:
      containers:
        - name: agent
          image: zhuquelab/aig-agent:latest
          env:
            - name: AIG_SERVER_URL
              value: http://aig-server:8088

```

### Enabling Prometheus Monitoring

Activate metrics collection to monitor queue depth and agent health:

```bash
AIG_ENABLE_METRICS=true ./ai-infra-guard webserver --server 0.0.0.0:8088

```

Access the `/metrics` endpoint to scrape data for autoscaling decisions and alerting on failed scan rates.

## Summary

- **Containerize** using `Dockerfile` for servers and `Dockerfile_Agent` for workers to ensure environment consistency across the cluster.
- **Scale horizontally** by running multiple server replicas behind a load balancer and increasing agent counts via [`common/websocket/task_manager.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/task_manager.go) task distribution.
- **Externalize state** to PostgreSQL or Redis using `AIG_DB_URL` to prevent data loss during rolling updates or crashes.
- **Monitor continuously** via the built-in Prometheus `/metrics` endpoint and configure autoscaling based on queue length.
- **Secure aggressively** by deploying only within private networks and using TLS termination at the load balancer, as the platform lacks native authentication.

## Frequently Asked Questions

### How many agent workers should I deploy per server instance?

Deploy **2-3 agents per server replica** as a baseline, then scale based on observed I/O wait times. Since [`common/websocket/task_manager.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/task_manager.go) distributes tasks concurrently, you can increase agent counts independently until you saturate network bandwidth or the database connection pool.

### Can I run AI-Infra-Guard server and agents on different networks?

Yes, provided the agents can reach the server's WebSocket endpoint via `AIG_SERVER_URL`. The [`common/websocket/websocket.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/websocket.go) implementation handles reconnection logic automatically, making it suitable for cross-VPC or hybrid cloud deployments using VPN peering or private link connections.

### What database permissions are required for production PostgreSQL?

The server requires standard CRUD permissions on tables storing task states and scan results. Configure a dedicated `aig` user with `CONNECT`, `CREATE`, `READ`, and `WRITE` privileges on the target database, but restrict `DROP` permissions to prevent accidental schema destruction during automated deployments.

### Is WebSocket sticky session support required for load balancing?

**No**, because the server in [`cmd/cli/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go) maintains task state in the external database rather than in-memory session stores. Agents reconnecting through a different server replica will resume work from the shared PostgreSQL or Redis store, allowing stateless load balancing across the server tier.