How to Scale AI-Infra-Guard in a Production Environment: 7-Step Deployment Guide

Deploy AI-Infra-Guard at scale by containerizing the Go-based server and lightweight agents, running multiple replicas behind a load balancer, and using external PostgreSQL or Redis for shared state persistence.

AI-Infra-Guard (AIG) is an open-source AI infrastructure scanner built by Tencent that combines a modular Go server with lightweight agent workers to audit machine learning environments. The repository (Tencent/AI-Infra-Guard) provides production-ready Docker images, WebSocket-based task distribution, and built-in Prometheus metrics, making it straightforward to scale from a single-node deployment to a high-availability cluster. This guide explains how to configure the core components—defined in [cmd/cli/main.go](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go) and [common/websocket/task_manager.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/common/websocket/task_manager.go)—for enterprise-grade workloads.

Architecture Overview

Before scaling, understand how the components interact in a distributed setup. The platform consists of four primary layers:

  • Server Layer (cmd/cli/main.go): The Go HTTP/WebSocket API that accepts scan requests, schedules tasks, and aggregates results.
  • Agent Workers (Dockerfile_Agent): Lightweight containers that register via WebSocket (common/websocket/websocket.go) and execute I/O-bound scans.
  • State Store: External PostgreSQL or Redis database configured via AIG_DB_URL to persist task states across server restarts.
  • Load Balancer: Distributes incoming HTTPS and WebSocket traffic across multiple server replicas.

The common/websocket/task_manager.go file implements the core queue logic, allowing agents to pick up pending jobs concurrently without blocking the main server process.

Step-by-Step Production Scaling

1. Containerize with Official Images

Start by using the pre-built images or building from the Dockerfiles in the repository. The Dockerfile creates the main AIG server containing the Go binary and web UI, while Dockerfile_Agent produces the minimal agent image used by MCP/Skill scanners.

Guarantee reproducible environments by pinning to specific image tags rather than latest. This isolation prevents dependency conflicts when running mixed versions across your cluster.

2. Implement Horizontal Scaling of Server Replicas

Deploy multiple instances of the server behind a load balancer such as Nginx, HAProxy, or a Kubernetes Service. Each replica in cmd/cli/main.go operates independently, accepting scan requests via the HTTP API while the WebSocket handler manages long-lived connections to agents.

Configure your load balancer to support sticky sessions for WebSocket connections or use a stateless round-robin for the REST API endpoints. This distribution prevents any single node from becoming a bottleneck during peak scanning periods.

3. Configure Shared State Persistence

Prevent data loss during pod restarts by externalizing the database. The server supports PostgreSQL, Redis, or mounted volumes via configuration flags defined in cmd/cli/main.go.

Set the AIG_DB_URL environment variable to point to your external store:

AIG_DB_URL=postgres://aig:password@db:5432/aig?sslmode=disable

This ensures that task states and scan results survive server crashes and allow newly spawned workers to resume pending jobs immediately.

4. Deploy Dedicated Agent Workers

Scale scan throughput by running multiple agent containers built from Dockerfile_Agent. These workers connect to the main server via WebSocket (common/websocket/websocket.go), register themselves automatically, and pull tasks from the distributed queue managed by common/websocket/task_manager.go.

Since scans are I/O-bound rather than CPU-bound, you can run many agent replicas on modest hardware. Each agent operates independently, so adding containers linearly increases throughput without overloading the server process.

5. Integrate the API-Checker Service

For model fingerprinting and relay checking capabilities, deploy the separate api-checker service located in services/api_checker. Configure the AIG_API_CHECKER_PYTHON environment variable as documented in the repository README.

Keeping this component isolated allows you to scale Python-based security checks independently from the Go server, preventing resource contention during intensive model analysis.

6. Enable Monitoring and Autoscaling

The Go server exposes Prometheus metrics at the /metrics endpoint when started with the --metrics flag or AIG_ENABLE_METRICS=true environment variable.

Configure your monitoring stack to scrape:

AIG_ENABLE_METRICS=true ./ai-infra-guard webserver --server 0.0.0.0:8088

Set up autoscaling rules based on CPU utilization, memory consumption, and queue depth metrics. This guarantees SLA compliance and automatic horizontal scaling during traffic spikes.

7. Secure the Deployment

AI-Infra-Guard has no built-in authentication. Always deploy inside a trusted network boundary (VPC or VPN) and restrict inbound ports to internal services only. Terminate TLS at the edge load balancer to encrypt WebSocket and HTTP traffic.

Never expose the scan APIs directly to the public internet, as this could allow unauthorized access to your infrastructure auditing capabilities.

Production Deployment Examples

Docker Compose with Replicas

The docker-compose.images.yml file provides a template for production deployment. Scale the server and agents independently using the deploy.replicas directive:

version: '3.8'
services:
  aig-server:
    image: zhuquelab/aig-server:latest
    restart: always
    environment:
      - AIG_DB_URL=postgres://aig:password@db:5432/aig?sslmode=disable
    ports:
      - "8088:8088"
    depends_on:
      - db
    deploy:
      mode: replicated
      replicas: 4
      resources:
        limits:
          cpus: "2"
          memory: 4G

  aig-agent:
    image: zhuquelab/aig-agent:latest
    restart: always
    environment:
      - AIG_SERVER_URL=http://aig-server:8088
    deploy:
      mode: replicated
      replicas: 8
      resources:
        limits:
          cpus: "1"
          memory: 2G

  db:
    image: postgres:15
    restart: always
    environment:
      POSTGRES_USER: aig
      POSTGRES_PASSWORD: password
      POSTGRES_DB: aig
    volumes:
      - db-data:/var/lib/postgresql/data

volumes:
  db-data:

Kubernetes Deployment

For orchestrated environments, separate the server and agent into distinct Deployments. The server handles API requests while agents scale based on scan backlog:

apiVersion: apps/v1
kind: Deployment
metadata:
  name: aig-server
spec:
  replicas: 3
  selector:
    matchLabels:
      app: aig-server
  template:
    metadata:
      labels:
        app: aig-server
    spec:
      containers:
        - name: server
          image: zhuquelab/aig-server:latest
          ports:
            - containerPort: 8088
          env:
            - name: AIG_DB_URL
              value: postgres://aig:password@postgres:5432/aig?sslmode=disable
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: aig-agent
spec:
  replicas: 6
  selector:
    matchLabels:
      app: aig-agent
  template:
    metadata:
      labels:
        app: aig-agent
    spec:
      containers:
        - name: agent
          image: zhuquelab/aig-agent:latest
          env:
            - name: AIG_SERVER_URL
              value: http://aig-server:8088

Enabling Prometheus Monitoring

Activate metrics collection to monitor queue depth and agent health:

AIG_ENABLE_METRICS=true ./ai-infra-guard webserver --server 0.0.0.0:8088

Access the /metrics endpoint to scrape data for autoscaling decisions and alerting on failed scan rates.

Summary

  • Containerize using Dockerfile for servers and Dockerfile_Agent for workers to ensure environment consistency across the cluster.
  • Scale horizontally by running multiple server replicas behind a load balancer and increasing agent counts via common/websocket/task_manager.go task distribution.
  • Externalize state to PostgreSQL or Redis using AIG_DB_URL to prevent data loss during rolling updates or crashes.
  • Monitor continuously via the built-in Prometheus /metrics endpoint and configure autoscaling based on queue length.
  • Secure aggressively by deploying only within private networks and using TLS termination at the load balancer, as the platform lacks native authentication.

Frequently Asked Questions

How many agent workers should I deploy per server instance?

Deploy 2-3 agents per server replica as a baseline, then scale based on observed I/O wait times. Since common/websocket/task_manager.go distributes tasks concurrently, you can increase agent counts independently until you saturate network bandwidth or the database connection pool.

Can I run AI-Infra-Guard server and agents on different networks?

Yes, provided the agents can reach the server's WebSocket endpoint via AIG_SERVER_URL. The common/websocket/websocket.go implementation handles reconnection logic automatically, making it suitable for cross-VPC or hybrid cloud deployments using VPN peering or private link connections.

What database permissions are required for production PostgreSQL?

The server requires standard CRUD permissions on tables storing task states and scan results. Configure a dedicated aig user with CONNECT, CREATE, READ, and WRITE privileges on the target database, but restrict DROP permissions to prevent accidental schema destruction during automated deployments.

Is WebSocket sticky session support required for load balancing?

No, because the server in cmd/cli/main.go maintains task state in the external database rather than in-memory session stores. Agents reconnecting through a different server replica will resume work from the shared PostgreSQL or Redis store, allowing stateless load balancing across the server tier.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →