Security Model of OpenHuman: Autonomy Tiers, Approval Gates, and Command Classification
OpenHuman protects users through a layered security model built into its Rust core, combining autonomy tiers that define baseline permissions, command classification that parses shell operations into risk categories, and approval gates that enforce explicit consent for elevated-risk actions.
OpenHuman is an AI agent runtime that executes shell commands and file operations on behalf of users. To prevent unauthorized data access or system modification, the platform implements a defense-in-depth architecture defined in the source code at tinyhumansai/openhuman. The security model ensures that every tool invocation passes through classification, tier verification, and optional user approval before execution.
Autonomy Tiers and Configuration
The foundation of OpenHuman’s security model is the autonomy tier system, which defines the default permission level for agent operations. The configuration schema lives in src/openhuman/config/schema/autonomy.rs and is exposed through the global Config object.
The AutonomyLevel enum defines three distinct tiers:
pub enum AutonomyLevel {
ReadOnly, // Only read-only operations are allowed
Supervised, // Write and network operations require approval gate clearance
Full, // All operations allowed unless explicitly blocked
}
Each tier supports granular controls configured via the autonomy block:
level: AutonomyLevel– Sets the base permission tierworkspace_only: bool– Confines actions to the workspace directory when enabledmax_actions_per_hour: u32– Implements rate limiting for tool usagerequire_approval_for_medium_risk: bool– Forces approval prompts for Write and Network commandsblock_high_risk_commands: bool– Automatically denies Destructive class operations
Users query or modify these settings at runtime through JSON-RPC methods implemented in src/openhuman/config/ops.rs:
openhuman.config_get_autonomy_settingsopenhuman.config_update_autonomy_settings
Command Classification System
Before any tool executes, OpenHuman parses the raw command string to determine its risk profile. The SecurityPolicy::classify_command function in src/openhuman/security/policy/command_checks.rs analyzes shell syntax, redirections, pipelines, and executable names to assign a CommandClass:
pub enum CommandClass {
Read, // e.g., cat file.txt
Write, // e.g., echo hi > out.txt or git operations
Network, // e.g., curl http://example.com
Install, // e.g., npm install (treated as Write for rate-limiting)
Destructive, // e.g., rm -rf /, shutdown, reboot
}
The classifier treats shell redirections (>, >>, |) as Write operations due to their side effects. It normalizes command names—for example, mapping git to the Write category—and flags hidden execution patterns like ${VAR} expansion combined with tee as potentially dangerous write operations.
Approval Gates and Enforcement Flow
When a command’s classification exceeds the current autonomy tier’s permissions, the approval gate determines the final execution path. The enforcement logic resides in src/openhuman/security/policy/enforcement.rs, specifically within SecurityPolicy::enforce_tool_operation.
The gate evaluates three possible outcomes:
- Allow – The command passes automatically if
autonomy.auto_approveis enabled and the risk level permits - Prompt – For medium-risk commands (Write or Network) when
require_approval_for_medium_riskis true, the gate generates anApprovalRequestcontaining the command description and timeout, suspending execution until the user responds via the UI - Block – High-risk Destructive commands are denied immediately when
block_high_risk_commandsis enabled
The decision logic is implemented in src/openhuman/security/approval/gate.rs, which bridges the autonomy configuration and command classification to produce the final verdict.
Runtime Guarantees and Sandboxing
OpenHuman implements multiple fail-safe mechanisms to ensure the security model cannot be bypassed:
- Fail-closed default – If classification or policy lookup fails, the system denies the operation rather than allowing potentially dangerous execution
- Rate limiting – The
max_actions_per_hourfield throttles tool usage across all autonomy tiers to prevent automated abuse - OS-level sandboxing – Even allowed commands execute within restricted environments using Landlock (Linux) or Firejail (cross-platform), enforced by
src/openhuman/security/landlock.rsandsrc/openhuman/security/firejail.rs - Comprehensive audit logging – Every decision (Allow, Prompt, or Block) is recorded to the event bus via
src/openhuman/security/audit.rsfor forensic analysis
Practical Configuration Examples
The following Rust examples demonstrate how to interact with the security model programmatically:
Query the current autonomy settings:
let resp = core_rpc_client
.call("openhuman.config_get_autonomy_settings", json!({}))
.await?;
println!("Current tier: {}", resp["level"]); // Outputs: "Supervised"
Update to full autonomy with auto-approval:
core_rpc_client
.call(
"openhuman.config_update_autonomy_settings",
json!({ "autonomy": { "level": "Full", "auto_approve": true } })
)
.await?;
Attempt a destructive command (will be blocked):
let tool_result = core_rpc_client
.call("openhuman.tools_shell", json!({ "command": "rm -rf /tmp/foo" }))
.await?;
assert_eq!(tool_result["status"], "error"); // Blocked as Destructive
Trigger a medium-risk approval request:
let tool_result = core_rpc_client
.call("openhuman.tools_shell", json!({ "command": "echo hi > out.txt" }))
.await?;
// Returns an ApprovalRequest ID; UI must approve before the write executes
Summary
- OpenHuman’s security model consists of autonomy tiers (ReadOnly, Supervised, Full) defined in
src/openhuman/config/schema/autonomy.rsthat establish baseline permissions - Command classification in
src/openhuman/security/policy/command_checks.rsparses shell commands into risk categories (Read, Write, Network, Install, Destructive) - Approval gates in
src/openhuman/security/approval/gate.rsenforce user consent for medium-risk operations and block high-risk commands automatically - The enforcement entry point
SecurityPolicy::enforce_tool_operationinsrc/openhuman/security/policy/enforcement.rsorchestrates all security decisions - Runtime guarantees include fail-closed defaults, rate limiting, OS-level sandboxing via Landlock/Firejail, and comprehensive audit logging
- Configuration is accessible through JSON-RPC methods
openhuman.config_get_autonomy_settingsandopenhuman.config_update_autonomy_settings
Frequently Asked Questions
How does OpenHuman classify complex shell commands with pipes and redirections?
According to the source code in src/openhuman/security/policy/command_checks.rs, the classifier treats any command containing shell redirections (>, >>) or pipes (|) as a Write operation, regardless of the primary executable. This conservative approach ensures that side effects like cat file | tee output.txt are properly flagged for approval gate review, preventing data exfiltration or modification through obfuscated command chains.
What happens if the approval gate cannot reach the user interface?
The approval gate implements a fail-closed policy. If the system cannot generate an ApprovalRequest or the user interface fails to respond within the timeout window, the operation is automatically denied. As implemented in src/openhuman/security/approval/gate.rs, this ensures that network interruptions or UI crashes never result in unauthorized high-risk command execution.
Can autonomy settings be changed dynamically without restarting the agent?
Yes. OpenHuman exposes runtime configuration updates through the openhuman.config_update_autonomy_settings JSON-RPC method defined in src/openhuman/config/ops.rs. Agents can transition between ReadOnly, Supervised, and Full tiers, or toggle auto_approve flags, without restarting the core runtime. All changes take effect immediately for subsequent tool invocations.
How does OpenHuman prevent catastrophic commands like rm -rf /?
Destructive commands are classified into the Destructive variant of the CommandClass enum. When block_high_risk_commands is enabled in the autonomy configuration, the approval gate in src/openhuman/security/approval/gate.rs automatically denies these operations before they reach the operating system. Additionally, even if classification fails, the sandbox layer (src/openhuman/security/landlock.rs) restricts filesystem access to the workspace directory, providing defense-in-depth against data loss.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →