# How AI-Infra-Guard Ensures Data Privacy for AI Workloads: A Layered Security Architecture

> AI-Infra-Guard secures AI workload data privacy with layered security: TLS encryption, process isolation, YAML scanning, and log redaction. Protect your AI.

- Repository: [Tencent/AI-Infra-Guard](https://github.com/tencent/AI-Infra-Guard)
- Tags: architecture
- Published: 2026-08-26

---

**AI-Infra-Guard protects AI workload confidentiality through TLS-encrypted WebSocket transport, container-level process isolation, privacy-specific YAML scanning rules, and automatic log redaction.**

Tencent's AI-Infra-Guard implements a defense-in-depth strategy to safeguard sensitive data throughout the AI infrastructure scanning lifecycle. The open-source project combines cryptographic protections with strict runtime isolation and targeted privacy detection signatures. This article examines the specific technical mechanisms implemented in the `Tencent/AI-Infra-Guard` repository that prevent data leakage during security assessments.

## Transport Security with TLS-Encrypted WebSockets

All communications between the central server and remote agents use **WebSocket over TLS (wss://)** to prevent eavesdropping and man-in-the-middle attacks. In [`cmd/agent/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/agent/main.go), the Go client establishes a TLS-encrypted connection before exchanging any payload data.

The implementation enforces **TLS 1.2** as the minimum version through the `tls.Config` structure. This ensures that data in transit remains confidential between distributed scanning agents and the control plane.

```go
// TLS-enabled WebSocket client implementation
import (
    "crypto/tls"
    "github.com/gorilla/websocket"
)

func newSecureWS(urlStr string) (*websocket.Conn, error) {
    dialer := websocket.Dialer{
        TLSClientConfig: &tls.Config{
            MinVersion: tls.VersionTLS12,
        },
    }
    return dialer.Dial(urlStr, nil)
}

```

## Container Isolation for Runtime Protection

Each scan executes inside a **dedicated Docker container** defined in [`docker-compose.yml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/docker-compose.yml), creating a sandboxed environment where inspected code cannot access the host filesystem or network directly. This containment strategy ensures that any accidental data leakage remains confined to the container's ephemeral filesystem, which is destroyed after the scan completes.

The [`docker-compose.yml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/docker-compose.yml) configuration isolates services with restricted network policies and non-persistent storage volumes, preventing scanned AI workloads from persisting sensitive data outside the temporary execution context.

## Privacy-Aware Rule Engine

The core scanner located in `internal/mcp` evaluates code and model artifacts against **privacy-specific signatures** stored as YAML vulnerability entries. These rules detect hard-coded API keys, insecure logging of user data, and exposure of invitation lists or email addresses.

For example, [`data/vuln_en/ragflow/CVE-2024-12869.yaml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/data/vuln_en/ragflow/CVE-2024-12869.yaml) defines a signature that identifies vulnerabilities potentially leaking email addresses or usernames, explicitly labeling such detections as "privacy breach" in the description field.

### Rule Validation and CI Enforcement

The [`cmd/yamlcheck/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/yamlcheck/main.go) utility validates all new rule files under `data/` before acceptance, ensuring privacy rules are syntactically correct and do not introduce false positives. The continuous integration pipeline executes `yamlcheck` on every pull request, guaranteeing that privacy detection logic remains accurate and up-to-date.

```bash

# Validate a custom privacy rule before committing

./yamlcheck data/custom_privacy.yaml

```

## Audit Logging with Automatic Redaction

All audit logs are written to **structured JSON** format with automatic redaction of fields marked as `privacy_sensitive`. The logic residing in [`pkg/log/redact.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/log/redact.go) strips any value whose key matches the privacy whitelist before persisting logs to disk.

This mechanism ensures that **operational logs never expose raw personal data**, even when scanning workloads that process sensitive user information. The redaction occurs at the logging package level, providing consistent privacy protection across all system components.

## Configurable Privacy Scanning Modes

Operators can enable a **privacy-only scan mode** using the `--privacy` CLI flag, which focuses solely on detecting personal-data leaks without performing full vulnerability analysis. The [`cmd/cli/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go) file parses this flag and loads only the privacy-related rule subset, reducing scan scope and minimizing data exposure risks.

```bash

# Run a privacy-focused scan against a target AI service

./ai-infra-guard scan -t http://127.0.0.1:8088 --privacy

# Add a custom privacy rule for email detection

cat > data/custom_privacy.yaml <<EOF
id: CUSTOM-PRIV-001
summary: "Prevent accidental logging of email addresses"
severity: low
description: |
  Detect any log statement that includes the field name `user_email`.
  This rule treats such occurrences as a privacy breach.
  reference: https://github.com/Tencent/AI-Infra-Guard/blob/main/data/custom_privacy.yaml
rule:
  pattern: "user_email"
  type: regex
EOF

```

## Summary

- **Data in transit** is protected via WebSocket over TLS (wss://) with minimum TLS 1.2 enforcement in [`cmd/agent/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/agent/main.go)
- **Data at rest** remains confined to ephemeral Docker containers defined in [`docker-compose.yml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/docker-compose.yml), preventing host filesystem access
- **Sensitive information detection** is performed by the `internal/mcp` engine using privacy-specific YAML signatures such as [`CVE-2024-12869.yaml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/CVE-2024-12869.yaml)
- **Operational logs** are sanitized by [`pkg/log/redact.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/log/redact.go) to remove privacy-sensitive fields before persistence
- **Scan scope limitation** is available through the `--privacy` flag in [`cmd/cli/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/cli/main.go) for targeted privacy assessments

## Frequently Asked Questions

### How does AI-Infra-Guard encrypt data in transit between agents and the server?

All agent-server communications use WebSocket connections over TLS (wss://) as implemented in [`cmd/agent/main.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/cmd/agent/main.go). The Go client configures a minimum TLS version of 1.2 before establishing the connection, ensuring end-to-end encryption that prevents network eavesdropping and man-in-the-middle attacks during scan operations.

### Can AI-Infra-Guard analyze sensitive AI workloads without exposing host system data?

Yes. The platform executes every scan inside isolated Docker containers configured in [`docker-compose.yml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/docker-compose.yml). This containerization ensures that the inspected code operates within an ephemeral filesystem that cannot access the host's network or persistent storage, containing any potential data leakage within the temporary sandbox.

### What types of privacy violations does the scanner detect?

The `internal/mcp` rule engine detects hard-coded authentication keys, insecure logging of personal identifiers, and exposure of user invitation lists. Specific signatures like [`data/vuln_en/ragflow/CVE-2024-12869.yaml`](https://github.com/Tencent/AI-Infra-Guard/blob/main/data/vuln_en/ragflow/CVE-2024-12869.yaml) target email address and username leaks, while the YAML-based rule system allows customization for organizational privacy requirements.

### How does the tool prevent sensitive data from appearing in audit logs?

The logging system in [`pkg/log/redact.go`](https://github.com/Tencent/AI-Infra-Guard/blob/main/pkg/log/redact.go) automatically redacts fields marked as `privacy_sensitive` before writing structured JSON logs. This redaction logic matches keys against a privacy whitelist and strips corresponding values, ensuring that audit trails contain operational metadata without exposing raw personal data from scanned workloads.