Security Model of OpenHuman: Autonomy Tiers, Approval Gates, and Command Classification

OpenHuman protects users through a layered security model built into its Rust core, combining autonomy tiers that define baseline permissions, command classification that parses shell operations into risk categories, and approval gates that enforce explicit consent for elevated-risk actions.

OpenHuman is an AI agent runtime that executes shell commands and file operations on behalf of users. To prevent unauthorized data access or system modification, the platform implements a defense-in-depth architecture defined in the source code at tinyhumansai/openhuman. The security model ensures that every tool invocation passes through classification, tier verification, and optional user approval before execution.

Autonomy Tiers and Configuration

The foundation of OpenHuman’s security model is the autonomy tier system, which defines the default permission level for agent operations. The configuration schema lives in src/openhuman/config/schema/autonomy.rs and is exposed through the global Config object.

The AutonomyLevel enum defines three distinct tiers:

pub enum AutonomyLevel {
    ReadOnly,   // Only read-only operations are allowed
    Supervised, // Write and network operations require approval gate clearance
    Full,       // All operations allowed unless explicitly blocked
}

Each tier supports granular controls configured via the autonomy block:

  • level: AutonomyLevel – Sets the base permission tier
  • workspace_only: bool – Confines actions to the workspace directory when enabled
  • max_actions_per_hour: u32 – Implements rate limiting for tool usage
  • require_approval_for_medium_risk: bool – Forces approval prompts for Write and Network commands
  • block_high_risk_commands: bool – Automatically denies Destructive class operations

Users query or modify these settings at runtime through JSON-RPC methods implemented in src/openhuman/config/ops.rs:

  • openhuman.config_get_autonomy_settings
  • openhuman.config_update_autonomy_settings

Command Classification System

Before any tool executes, OpenHuman parses the raw command string to determine its risk profile. The SecurityPolicy::classify_command function in src/openhuman/security/policy/command_checks.rs analyzes shell syntax, redirections, pipelines, and executable names to assign a CommandClass:

pub enum CommandClass {
    Read,        // e.g., cat file.txt
    Write,       // e.g., echo hi > out.txt or git operations
    Network,     // e.g., curl http://example.com
    Install,     // e.g., npm install (treated as Write for rate-limiting)
    Destructive, // e.g., rm -rf /, shutdown, reboot
}

The classifier treats shell redirections (>, >>, |) as Write operations due to their side effects. It normalizes command names—for example, mapping git to the Write category—and flags hidden execution patterns like ${VAR} expansion combined with tee as potentially dangerous write operations.

Approval Gates and Enforcement Flow

When a command’s classification exceeds the current autonomy tier’s permissions, the approval gate determines the final execution path. The enforcement logic resides in src/openhuman/security/policy/enforcement.rs, specifically within SecurityPolicy::enforce_tool_operation.

The gate evaluates three possible outcomes:

  1. Allow – The command passes automatically if autonomy.auto_approve is enabled and the risk level permits
  2. Prompt – For medium-risk commands (Write or Network) when require_approval_for_medium_risk is true, the gate generates an ApprovalRequest containing the command description and timeout, suspending execution until the user responds via the UI
  3. Block – High-risk Destructive commands are denied immediately when block_high_risk_commands is enabled

The decision logic is implemented in src/openhuman/security/approval/gate.rs, which bridges the autonomy configuration and command classification to produce the final verdict.

Runtime Guarantees and Sandboxing

OpenHuman implements multiple fail-safe mechanisms to ensure the security model cannot be bypassed:

  • Fail-closed default – If classification or policy lookup fails, the system denies the operation rather than allowing potentially dangerous execution
  • Rate limiting – The max_actions_per_hour field throttles tool usage across all autonomy tiers to prevent automated abuse
  • OS-level sandboxing – Even allowed commands execute within restricted environments using Landlock (Linux) or Firejail (cross-platform), enforced by src/openhuman/security/landlock.rs and src/openhuman/security/firejail.rs
  • Comprehensive audit logging – Every decision (Allow, Prompt, or Block) is recorded to the event bus via src/openhuman/security/audit.rs for forensic analysis

Practical Configuration Examples

The following Rust examples demonstrate how to interact with the security model programmatically:

Query the current autonomy settings:

let resp = core_rpc_client
    .call("openhuman.config_get_autonomy_settings", json!({}))
    .await?;
println!("Current tier: {}", resp["level"]); // Outputs: "Supervised"

Update to full autonomy with auto-approval:

core_rpc_client
    .call(
        "openhuman.config_update_autonomy_settings",
        json!({ "autonomy": { "level": "Full", "auto_approve": true } })
    )
    .await?;

Attempt a destructive command (will be blocked):

let tool_result = core_rpc_client
    .call("openhuman.tools_shell", json!({ "command": "rm -rf /tmp/foo" }))
    .await?;
assert_eq!(tool_result["status"], "error"); // Blocked as Destructive

Trigger a medium-risk approval request:

let tool_result = core_rpc_client
    .call("openhuman.tools_shell", json!({ "command": "echo hi > out.txt" }))
    .await?;
// Returns an ApprovalRequest ID; UI must approve before the write executes

Summary

  • OpenHuman’s security model consists of autonomy tiers (ReadOnly, Supervised, Full) defined in src/openhuman/config/schema/autonomy.rs that establish baseline permissions
  • Command classification in src/openhuman/security/policy/command_checks.rs parses shell commands into risk categories (Read, Write, Network, Install, Destructive)
  • Approval gates in src/openhuman/security/approval/gate.rs enforce user consent for medium-risk operations and block high-risk commands automatically
  • The enforcement entry point SecurityPolicy::enforce_tool_operation in src/openhuman/security/policy/enforcement.rs orchestrates all security decisions
  • Runtime guarantees include fail-closed defaults, rate limiting, OS-level sandboxing via Landlock/Firejail, and comprehensive audit logging
  • Configuration is accessible through JSON-RPC methods openhuman.config_get_autonomy_settings and openhuman.config_update_autonomy_settings

Frequently Asked Questions

How does OpenHuman classify complex shell commands with pipes and redirections?

According to the source code in src/openhuman/security/policy/command_checks.rs, the classifier treats any command containing shell redirections (>, >>) or pipes (|) as a Write operation, regardless of the primary executable. This conservative approach ensures that side effects like cat file | tee output.txt are properly flagged for approval gate review, preventing data exfiltration or modification through obfuscated command chains.

What happens if the approval gate cannot reach the user interface?

The approval gate implements a fail-closed policy. If the system cannot generate an ApprovalRequest or the user interface fails to respond within the timeout window, the operation is automatically denied. As implemented in src/openhuman/security/approval/gate.rs, this ensures that network interruptions or UI crashes never result in unauthorized high-risk command execution.

Can autonomy settings be changed dynamically without restarting the agent?

Yes. OpenHuman exposes runtime configuration updates through the openhuman.config_update_autonomy_settings JSON-RPC method defined in src/openhuman/config/ops.rs. Agents can transition between ReadOnly, Supervised, and Full tiers, or toggle auto_approve flags, without restarting the core runtime. All changes take effect immediately for subsequent tool invocations.

How does OpenHuman prevent catastrophic commands like rm -rf /?

Destructive commands are classified into the Destructive variant of the CommandClass enum. When block_high_risk_commands is enabled in the autonomy configuration, the approval gate in src/openhuman/security/approval/gate.rs automatically denies these operations before they reach the operating system. Additionally, even if classification fails, the sandbox layer (src/openhuman/security/landlock.rs) restricts filesystem access to the workspace directory, providing defense-in-depth against data loss.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →