How the plan_gate Module Controls Agent Permissions in orx

The plan_gate module implements a Pre-Tool-Use hook that intercepts agent commands in Claude's --permission-mode plan, applying a strict allow-list policy to automatically approve read-only operations while gating potentially destructive actions for user approval.

The plan_gate module in the alphaXiv/OpenResearch repository serves as the security gateway between Claude's plan mode and the orx CLI. When running with --permission-mode plan, Claude delegates permission decisions to this Rust-based classifier, which examines every tool invocation before execution to determine whether it can run unsupervised or requires explicit user consent.

The Pre-Tool-Use Hook Architecture

The module operates as a bridge between the LLM agent and the operating system, receiving JSON payloads via stdin and returning permission decisions.

Hook Entry Point

In src/commands/plan_gate.rs, the executable reads the JSON payload from Claude on stdin and forwards it to the core classifier. This CLI entry point handles the I/O boundaries before delegating to plan_gate_decide, which is defined in src/local/harness/plan_gate.rs and re-exported through src/local/harness/mod.rs as pub use plan_gate::decide as plan_gate_decide.

Decision Logic

The decide(payload) function examines the tool_name field to categorize the request:

  • ExitPlanMode – Returns an ask decision, forcing explicit user approval before exiting plan mode (lines 99-110 in src/local/harness/plan_gate.rs).
  • Bash – Extracts the command string from /tool_input/command and routes it through command_is_readonly for safety analysis (lines 112-124).
  • Other Tools – Defers to default plan-mode gating by returning None, triggering the standard user prompt.

Classifying Read-Only vs. Gated Commands

The command_is_readonly function implements a strict allow-list policy that treats all commands as potentially dangerous unless explicitly proven safe.

Dangerous Metacharacter Detection

Before parsing command structure, the classifier rejects any input containing dangerous metacharacters. The validation fails immediately if the command contains backticks, <, $(, or newlines (lines 64-69 in src/local/harness/plan_gate.rs), preventing command substitution and redirection attacks that could bypass the read-only restriction.

Pipeline Validation

After sanitization, the function splits the command on && and ; delimiters, requiring every segment to pass the is_readonly_segment check (lines 71-77). Each segment may be a simple shell no-op, a single pipeline, or a sequence of pipelines. Pipelines are accepted only when the first stage is a read-only producer and all subsequent stages are pure consumers (lines 78-102).

Allow-List Categories for Safe Commands

The is_readonly_producer function maintains explicit lists of permissible commands, categorizing them by their proven safety profiles.

Glue Commands

Simple shell utilities like echo, true, and : (the null command) are always allowed (lines 34-36). These serve as harmless separators or status checks within complex command chains.

orx Read-Only Verbs

orx invocations are permitted when the top-level verb appears in WHOLE_VERB_READS (including runs, logs, and discover), or when sub-commands match known read-only patterns like orx project view or orx exp desc without write flags such as --set or --stdin (lines 35-73 and 74-84).

Git Read-Only Operations

Git commands are validated against the GIT_WHOLE_VERB_READS list, ensuring only inspection operations like status, log, or diff pass through. The classifier also scans for write-affecting arguments such as --output or -O, rejecting any git invocation that might modify repository state (lines 54-60 and 93-101).

Pure Consumer Utilities

Downstream pipeline stages (after the initial producer) are restricted to pure consumers like head, grep, wc, sort, and uniq (lines 45-52). These utilities can process data but cannot initiate write operations or execute subcommands.

Safety Mechanisms and Tokenization

Each pipeline stage undergoes tokenization via the stage_tokens function to detect redirection attempts.

Stage Tokenization

The tokenizer processes shell syntax to identify file descriptor redirections. Tokens like 2>&1 are harmlessly dropped as stderr-to-stdout redirects, while any token containing > (file overwrite) or stray & (background execution) causes immediate rejection of the entire stage (lines 103-119 in src/local/harness/plan_gate.rs).

Security Guarantees

By maintaining a hand-maintained allow-list rather than a block-list, the module ensures that any newly added write-capable verb in orx or git is automatically gated until explicitly added to the safe list. This architectural choice preserves the security invariant that a read-only prefix can never smuggle a write operation through command chaining or shell escapes.

Implementation Example

The following Rust example demonstrates how the hook receives and processes a JSON payload from Claude:

// The hook receives a JSON payload from Claude via stdin
let payload = json!({
    "tool_name": "Bash",
    "tool_input": { "command": "orx runs && orx logs r-1" }
});

if let Some(decision) = plan_gate_decide(&payload) {
    // Prints {"permissionDecision":"allow"} for read-only chains
    println!("{}", decision);
}

You can also invoke the classifier directly for testing command safety:

// Direct invocation for unit testing permission logic
assert!(command_is_readonly("orx runs; echo ====; orx logs r-1"));
assert!(!command_is_readonly("orx exp run e-1")); // Write verb → gated

Summary

  • The plan_gate module acts as a Pre-Tool-Use hook for Claude's --permission-mode plan, intercepting commands before execution.
  • ExitPlanMode always triggers user approval to prevent accidental plan exit, while Bash commands undergo deep inspection via command_is_readonly.
  • The classifier uses a strict allow-list policy, rejecting any command containing dangerous metacharacters or undefined verbs.
  • Safe commands include specific orx read-only verbs (runs, logs, discover), git inspection commands, and pure consumer utilities like grep and wc.
  • All pipeline stages are tokenized to detect file redirections and background execution attempts.
  • The hand-maintained allow-list ensures new write-capable features are automatically gated until explicitly approved.

Frequently Asked Questions

What triggers the plan_gate to gate a command?

Any command that fails the command_is_readonly check is gated for user approval. This includes Bash commands containing dangerous metacharacters (backticks, $(), <, newlines), unrecognized orx verbs not in WHOLE_VERB_READS, git commands outside the read-only list, or any pipeline containing file redirection operators like >.

How does plan_gate handle Bash commands with multiple segments?

The command_is_readonly function splits commands on && and ; delimiters, then validates every segment independently using is_readonly_segment. If any segment performs a write operation or uses an unrecognized command, the entire command chain returns None, forcing the default plan-mode gate to prompt the user.

Why does ExitPlanMode always require user approval?

The decide function explicitly checks for the ExitPlanMode tool name and returns an ask decision regardless of context. This prevents the agent from silently exiting plan mode and potentially executing unsupervised actions, ensuring the user maintains explicit control over when the permission gate disengages.

Where is the plan_gate_decide function exposed?

The core decision logic resides in src/local/harness/plan_gate.rs, but it is publicly exposed through src/local/harness/mod.rs using pub use plan_gate::decide as plan_gate_decide. The CLI entry point in src/commands/plan_gate.rs consumes this function to bridge stdin input from Claude with the classification logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →