How to Integrate AI-Infra-Guard with Existing ML Pipelines: Complete API Guide
AI-Infra-Guard integrates with existing ML pipelines through its REST API and Python helper modules, allowing orchestration tools like Airflow or Kubeflow to trigger security scans via HTTP endpoints.
Tencent's AI-Infra-Guard (A.I.G) is a stand-alone Go-based web service designed for AI infrastructure security scanning. Because it exposes a language-agnostic REST API, you can embed security audits into any stage of your machine learning workflow—whether during preprocessing, model training, or post-deployment validation—without rewriting your existing pipeline logic.
Architecture Overview
The integration capability stems from three core components in the repository:
cmd/cli/main.go– The entry point that hosts the HTTP and WebSocket server on port 8088.common/websocket/api.go– Defines the REST endpoint groups for file upload, task creation, and status retrieval.common/websocket/task_manager.go– Handles task distribution, assigning jobs to available Agents and persisting state to the database.
This architecture ensures that your ML pipeline only needs to speak HTTP to offload security scanning work to AI-Infra-Guard.
Integration Methods
You can integrate AI-Infra-Guard using either raw HTTP calls or the official Python wrappers.
REST API Integration
The service exposes four primary endpoints for pipeline integration:
/api/v1/app/taskapi/upload– Optional file upload for source code archives (used withmcp_scan)./api/v1/app/taskapi/tasks– Creates a scan task with a JSON payload specifying the scan type.- **
/api/v1/app/taskapi/status/{id}``** – Polls task progress (pending,running,completed, orfailed`). - `/api/v1/app/taskapi/result/{id}`` – Retrieves the final vulnerability report.
Supported scan types include agent_scan, mcp_scan, ai_infra_scan, and model_redteam_report.
Python Package Integration
For Python-centric workflows, install the helper modules via pip:
aig-agent-scanaig-mcp-scanaig-skill-scan
These packages function as CLI tools that wrap the same REST endpoints, allowing you to embed a single shell command inside your requirements.txt-driven environment without managing HTTP logic manually.
Step-by-Step Implementation
Follow this sequence to embed AI-Infra-Guard into your ML pipeline.
1. Start the Server
Deploy the service using Docker or the binary distribution as described in the quick-start section of README.md. The server must be accessible from your pipeline workers at the configured host and port (default 8088).
2. Upload Assets (Optional)
If scanning MCP server source code, upload the archive first:
import requests
base = "http://localhost:8088"
upload_url = f"{base}/api/v1/app/taskapi/upload"
with open("mcp_code.zip", "rb") as f:
resp = requests.post(upload_url, files={"file": f})
file_url = resp.json()["data"]["fileUrl"]
3. Create a Scan Task
Submit a JSON payload defining the scan configuration. The following example initiates an Agent scan with an inline YAML configuration:
import requests
import json
task_url = f"{base}/api/v1/app/taskapi/tasks"
yaml_cfg = """
provider: dify
base_url: https://your-dify-instance.example.com
api_key: app-your-dify-api-key
"""
payload = {
"type": "agent_scan",
"content": {
"agent_config": yaml_cfg,
"eval_model": {
"model": "gpt-4",
"token": "sk-your-api-key",
"base_url": "https://api.openai.com/v1"
},
"language": "en",
"prompt": "Focus on privilege escalation and data leakage"
}
}
resp = requests.post(task_url, json=payload)
session_id = resp.json()["data"]["session_id"]
print("Created task, session ID:", session_id)
4. Poll for Status and Retrieve Results
Poll the status endpoint until completion, then fetch the report:
import time
def wait_until_done(sid):
while True:
r = requests.get(f"{base}/api/v1/app/taskapi/status/{sid}")
data = r.json()["data"]
print("Status:", data["status"])
if data["status"] in ("completed", "failed"):
return data["status"]
time.sleep(10)
def fetch_result(sid):
r = requests.get(f"{base}/api/v1/app/taskapi/result/{sid}")
return r.json()["data"]
if wait_until_done(session_id) == "completed":
result = fetch_result(session_id)
print(json.dumps(result, indent=2, ensure_ascii=False))
Code Examples for Orchestration Tools
Shell One-Liner for CI/CD
Execute a complete MCP scan from any shell-based pipeline:
# Upload source archive
FILE_URL=$(curl -s -X POST http://localhost:8088/api/v1/app/taskapi/upload \
-F "file=@my_mcp_code.zip" | jq -r .data.fileUrl)
# Create scan task
SESSION=$(curl -s -X POST http://localhost:8088/api/v1/app/taskapi/tasks \
-H "Content-Type: application/json" \
-d "{\"type\":\"mcp_scan\",\"content\":{\"prompt\":\"Scan this MCP server\",\"model\":{\"model\":\"gpt-4\",\"token\":\"sk-your-key\",\"base_url\":\"https://api.openai.com/v1\"},\"thread\":4,\"language\":\"en\",\"attachments\":\"$FILE_URL\"}}" \
| jq -r .data.session_id)
# Poll until finished
while :; do
STATUS=$(curl -s http://localhost:8088/api/v1/app/taskapi/status/$SESSION | jq -r .data.status)
echo "Current status: $STATUS"
[[ $STATUS == "completed" ]] && break
sleep 10
done
# Retrieve report
curl http://localhost:8088/api/v1/app/taskapi/result/$SESSION | jq .
Apache Airflow DAG Example
Embed the scan as a PythonOperator task within your Airflow DAG:
from airflow import DAG
from airflow.operators.python import PythonOperator
from datetime import datetime
import requests
import time
import json
def run_aig_task(**kwargs):
base = "http://ai-guard:8088"
# Create AI-Infra scan task
payload = {
"type": "ai_infra_scan",
"content": {
"target": ["http://my-llm-service:8000"],
"headers": {"Authorization": "Bearer $TOKEN"},
"model": {
"model": "gpt-4",
"token": "sk-key",
"base_url": "https://api.openai.com/v1"
}
}
}
resp = requests.post(f"{base}/api/v1/app/taskapi/tasks", json=payload)
sid = resp.json()["data"]["session_id"]
# Wait for completion
while True:
st = requests.get(f"{base}/api/v1/app/taskapi/status/{sid}").json()["data"]["status"]
if st == "completed":
break
time.sleep(5)
result = requests.get(f"{base}/api/v1/app/taskapi/result/{sid}").json()["data"]
return result
with DAG("aig_integration", start_date=datetime(2024, 1, 1), schedule_interval=None) as dag:
security_scan = PythonOperator(
task_id="run_aig_security_scan",
python_callable=run_aig_task
)
Summary
- AI-Infra-Guard operates as a stand-alone Go service with a well-defined REST API, making it compatible with any ML orchestration tool.
- Key files for integration include
cmd/cli/main.go(server entry),common/websocket/api.go(endpoint definitions), andcommon/websocket/task_manager.go(task orchestration). - Primary endpoints are
/api/v1/app/taskapi/upload,/tasks,/status/{id}, and/result/{id}. - Python packages (
aig-agent-scan,aig-mcp-scan,aig-skill-scan) provide CLI wrappers for Python-centric pipelines. - Scan types cover Agent evaluation, MCP server auditing, AI infrastructure scanning, and model red-teaming.
Frequently Asked Questions
What orchestration tools work with AI-Infra-Guard?
Any tool capable of making HTTP requests or executing shell commands works with AI-Infra-Guard. Popular ML orchestration platforms like Apache Airflow, Kubeflow Pipelines, Prefect, and MLflow can trigger scans using standard PythonOperator or BashOperator tasks. The service's language-agnostic REST API ensures compatibility regardless of your pipeline's underlying framework.
Do I need to modify my existing ML code to use AI-Infra-Guard?
No code modification is required. Because AI-Infra-Guard runs as a separate service, your pipeline interacts with it purely through network calls. You can add security scanning as a discrete step—similar to how you might call a model training API—without changing your preprocessing or inference logic. For Python environments, installing the helper packages adds CLI commands that handle the HTTP communication transparently.
Can AI-Infra-Guard scan assets stored in my pipeline's object storage?
Yes. The /api/v1/app/taskapi/upload endpoint accepts file uploads, allowing you to stream artifacts directly from your pipeline's storage (S3, GCS, or local volumes) into AI-Infra-Guard for analysis. Alternatively, if your infrastructure supports pre-signed URLs, you can pass accessible URLs in the task payload's attachments field rather than uploading raw bytes.
How do I handle authentication between my pipeline and AI-Infra-Guard?
As of the current implementation, the open-source version exposes the REST API without built-in authentication headers. For production ML pipelines, deploy AI-Infra-Guard behind a reverse proxy (nginx, Envoy, or an API gateway) that handles TLS termination and bearer token validation. Your pipeline then includes the authentication headers in requests to the proxy endpoint, securing the communication channel while maintaining the simple HTTP interface.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →