How Grox Classifies Posts Using LLM Prompts: Inside the Safety Screening Architecture

Grox converts raw social-media posts into structured safety decisions through a five-step pipeline that uses VisionSampler for model communication, carefully crafted Conversation prompts, and Pydantic validation to output typed PostSafetyScreenResult objects.

The xai-org/x-algorithm repository implements a prompt-centric classification system designed to evaluate content safety at scale. Understanding how Grox classify posts using LLM prompts reveals an architecture built around modularity, type safety, and clear separation between prompt engineering, rendering, and inference logic.

The Five-Step Classification Pipeline

The PostSafetyDeluxeClassifier orchestrates the entire workflow, transforming raw post data into actionable safety decisions through discrete, reversible stages.

Step 1: Initializing the VisionSampler

Every classification begins with instantiating a sampler capable of handling mixed text-and-image content. The VisionSampler class in grox/libs/grok_sampler/vision_sampler.py wraps the target LLM (such as grok-4-1-fast-hedgehog-critical) and manages the serialization of requests into the SampleTextRequest format expected by the underlying service.

This abstraction allows the pipeline to switch between model providers without modifying classification logic, as the sampler encapsulates model-specific configurations including temperature, nucleus probability settings, and JSON schema hints.

Step 2: Building the Conversation Prompt

The classifier constructs a Conversation object through the build_convo method defined in grox/flows/upa/classifier_post_safety_screen_deluxe.py. This process follows a strict three-message pattern:

  1. A system prompt containing the policy definition (sourced from grox/flows/upa/prompts.py) establishes the LLM's role and safety criteria.
  2. A user message containing the rendered user profile and post content, generated by UserRenderer (grox/core/lm/user.py) and PostRenderer (grox/core/lm/post.py).
  3. An empty assistant message signaling to the LLM where its generated response should begin.

This structured conversation format ensures consistent LLM behavior across different post types and policy checks.

Step 3: Executing the LLM Request

The assembled conversation undergoes "interleaving"—flattening into a list of strings and images—before transmission. The VisionSampler.sample method constructs the final SampleTextRequest via _get_sample_request (implemented in grox/libs/grok_sampler/llm.py), handling the HTTP or GRPC transport to the LLM service.

This stage manages all model-specific parameters, including sampling strategies and optional structured output constraints, keeping the classifier agnostic to transport protocols.

Step 4: Parsing Structured JSON Output

Upon receiving the LLM's free-form text response, the _parse method in grox/flows/upa/classifier_post_safety_screen_deluxe.py extracts content using compiled regex patterns. Grox expects responses containing human-readable reasoning followed by a <json>...</json> block.

The regex isolates the JSON payload, which is then validated against the PostSafetyScreenResult Pydantic model. This validation step guarantees type safety and eliminates ad-hoc string parsing in downstream components.

Step 5: Returning Typed Safety Results

The final output is a strongly-typed PostSafetyScreenResult object defined in grox/flows/upa/state_post_safety.py. This structured result contains the safety decision, confidence scores, and policy violations, consumable immediately by moderation engines and safety pipelines without further transformation.

Architectural Design Patterns

The Grox codebase demonstrates several sophisticated patterns that enable rapid iteration on classification tasks while maintaining system reliability.

Prompt-Centric Modularity

Every classifier acts as a thin wrapper defining how to ask the LLM, while heavy-lifting components like tokenization and transport live in the LiteLLM base class. This design makes it trivial to swap models or add new modalities (such as vision capabilities) by changing sampler configurations in grox/config/config.py rather than rewriting classification logic.

Type Safety with Pydantic Models

The rigid <json> block parsing into Pydantic models ensures that downstream systems receive well-formed objects. The PostSafetyScreenResult model serves as a contract between the LLM interpretation layer and the safety engine, preventing runtime errors from malformed outputs.

Separation of Concerns

The architecture cleanly divides responsibilities across three layers:

  • Prompt Generation: Centralized in grox/flows/upa/prompts.py, maintaining policy definitions separately from execution logic.
  • Content Rendering: UserRenderer and PostRenderer handle the transformation of internal data structures into LLM-ready textual representations.
  • Inference Management: VisionSampler and its base classes manage API communication, retry logic, and multimodal support.

Implementation Examples

The following examples demonstrate practical usage of the classification pipeline, from high-level classifier invocation to low-level sampler interaction.

Running the Deluxe Safety Classifier

from grox.flows.upa.classifier_post_safety_screen_deluxe import PostSafetyDeluxeClassifier
import asyncio

async def run_safety_check(post):
    """
    Classify a single post using the deluxe safety screen.
    `post` must be an instance of grox.core.data_loaders.data_types.Post
    """
    classifier = PostSafetyDeluxeClassifier()
    result = await classifier.classify(post)
    print("Safety decision:", result)
    return result

# Execute in an async context

# asyncio.run(run_safety_check(my_post))

Direct VisionSampler Usage

from grox.libs.grok_sampler.vision_sampler import VisionSampler
from grok_sampler.config import GrokModelConfig
import grox.config.config as grox_config

# Load model configuration from global Grox config

model_cfg = GrokModelConfig(
    **grox_config.get_model("grok-4-1-fast-hedgehog-critical").model_dump()
)
sampler = VisionSampler(model_cfg)

# Build a simple conversation

conversation = [
    "You are a safety analyst. Decide if the following tweet violates policy.",
    "User: @alice\nTweet: I love cats."
]

# Execute sampling

response = await sampler.sample(conversation, conversation_id="demo123")
print(response)

Summary

Grox implements a robust, extensible system to classify posts using LLM prompts through the following architectural decisions:

  • VisionSampler abstraction in grox/libs/grok_sampler/vision_sampler.py handles multimodal LLM communication and request serialization.
  • Conversation-based prompting via PostSafetyDeluxeClassifier.build_convo ensures consistent message formatting with system prompts, rendered content, and assistant delimiters.
  • Structured output parsing using regex extraction and Pydantic validation via PostSafetyScreenResult guarantees type-safe downstream consumption.
  • Config-driven model selection enables per-task tuning without code changes, reading from grox/config/config.py.
  • Reusable pipeline pattern extends beyond safety screening to banger detection, PTOS policy checks, and spam scoring workflows.

Frequently Asked Questions

How does Grox handle multimodal posts containing images in the classification pipeline?

The VisionSampler class specifically manages mixed text-and-image content by serializing inputs into SampleTextRequest objects that the LLM service understands. When processing posts with media, the sampler interleaves image data with text content before transmission, allowing the same PostSafetyDeluxeClassifier logic to handle both text-only and multimodal safety evaluations without modification.

What makes the Grox classification pipeline extensible to new safety policies?

The pipeline uses a prompt-centric design where new classifiers only need to define unique system prompts in grox/flows/upa/prompts.py and corresponding Pydantic result models. The heavy lifting of request construction, LLM communication, and JSON parsing remains in the reusable VisionSampler and LiteLLM base classes, allowing developers to implement new classification flows by changing configuration rather than infrastructure code.

How does Grox ensure reliable parsing of LLM outputs into structured data?

Grox requires LLM responses to include both human-readable reasoning and a strict <json>...</json> XML block. The _parse method in classifier_post_safety_screen_deluxe.py uses compiled regex patterns to extract this block, then validates the content against the PostSafetyScreenResult Pydantic model defined in grox/flows/upa/state_post_safety.py. This two-step extraction and validation process prevents malformed outputs from propagating into production safety systems.

Where are the specific LLM models configured in the Grox codebase?

Model configurations reside in grox/config/config.py, accessed through grox_config.get_model(). This centralizes model selection (such as grok-4-1-fast-hedgehog-critical or Gemma variants), temperature settings, and API credentials. The VisionSampler instantiates using GrokModelConfig objects derived from these global settings, enabling per-task model tuning—such as using faster models for initial screens and larger models for appeals—without modifying classifier implementations.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →