Safety Mechanisms and Confirmation Steps for AI File Deletion Tools

AI assistants implement a defense-in-depth strategy for file deletion that combines schema-level path validation, mandatory user confirmation via UI elements like display_draft, tool-only policies forbidding shell commands, and sandbox-based undo capabilities to prevent accidental or malicious data loss.

The x1xhlol/system-prompts-and-models-of-ai-tools repository reveals how modern AI coding assistants protect users from catastrophic data loss when invoking delete_file or equivalent tools. These safety mechanisms and confirmation steps are embedded across multiple layers—from JSON tool schemas to system prompts—ensuring that file removal requires explicit, informed user consent and operates within strict sandbox constraints.

Schema-Level Safety Mechanisms in Tool Definitions

Graceful Failure and Path Validation

The foundation of safe file deletion begins in the tool schema itself. In Same.dev/Tools.json, the delete_file tool specification explicitly states that the operation will fail gracefully if the target file does not exist, is protected, or if the operation is rejected. This prevents the assistant from crashing the workflow or generating errors when attempting to delete non-existent resources.

Path traversal attacks are mitigated through strict relative path requirements. The schema mandates that the tool expects a relative path rooted at the workspace directory, effectively blocking attempts to access files outside the repository via patterns like ../../secret.txt.

Existence Verification and Single-File Granularity

The Trae/Builder Tools.json specification adds a critical constraint: the assistant must verify the file exists before attempting deletion. The prompt explicitly states, "you MUST make sure the files is exist before deleting," preventing errors on missing targets.

While some implementations allow multiple files per call, the design emphasizes single-file granularity or explicit multi-file enumeration. This prevents bulk accidental deletions where a single mistaken command could wipe out entire directories.

Mandatory User Confirmation Workflows

Before any file removal occurs, the assistant must obtain explicit user confirmation. The v0 Prompts and Tools/Prompt.txt file outlines a structured workflow using task tracking elements like taskNameActive='Deleting file'. The assistant must:

  1. Clearly state which file it intends to delete
  2. Ask the user for explicit confirmation
  3. Wait for a positive response before invoking the tool

This three-step process ensures the human maintains veto power over every destructive operation.

UI-Based Transparency with display_draft

To prevent confirmation messages from being buried in text streams, Poke/Poke_p3.txt mandates the use of the display_draft tool. When confirming deletion, the assistant must wrap output using this UI helper, presenting the request in a human-readable format that clearly identifies the target resource.

This requirement ensures transparency by making the deletion request visually distinct and unambiguous, reducing the risk of users accidentally approving destructive actions.

Operational Constraints and Policy Enforcement

Tool-Only Deletion Policies

A critical safety layer involves forbidding alternative deletion methods. The Augment Code/claude-4-sonnet-tools.json specification explicitly states that assistants must only use the dedicated delete_file tool for removing files. The policy strictly prohibits using shell commands like rm -rf or launch-process tools for file removal.

This constraint ensures that all deletions are tracked by the AI runtime, subject to the schema's safety checks, and potentially reversible through the sandbox's recording mechanisms.

Protection of Critical Files

System prompts add context-aware constraints to prevent catastrophic project damage. The Orchids.app/System Prompt.txt explicitly instructs the assistant: "Do not delete directories or critical configuration files."

This high-level policy prevents the AI from removing essential project infrastructure—such as build scripts, configuration files, or entire directory trees—even when the user might have technically granted broader permissions.

Recovery and Undo Capabilities

Beyond prevention, the repository implements post-deletion safety nets. The Orchids.app/System Prompt.txt references an E2B sandbox implementation where the delete_file tool operates within a controlled environment.

The runtime records a diff of the deletion, enabling the assistant to restore the file if the user changes their mind. By issuing a subsequent create_file call with the recorded content, the assistant can effectively undo the deletion, providing a recovery path even after the operation completes.

Summary

  • Schema-level protections in files like Same.dev/Tools.json enforce graceful failure, relative path constraints, and prevention of directory traversal attacks.
  • Mandatory confirmation workflows require explicit user consent through structured prompts and UI elements like display_draft before any deletion occurs.
  • Tool-only policies strictly prohibit shell-based deletions, ensuring all file removals are tracked and subject to safety checks as specified in Augment Code/claude-4-sonnet-tools.json.
  • Existence verification requirements in Trae/Builder Tools.json prevent errors by ensuring files are present before deletion attempts.
  • Recovery capabilities via sandbox diff recording allow assistants to restore deleted files if users change their minds.

Frequently Asked Questions

What prevents an AI assistant from deleting files outside the workspace?

The delete_file tool schema in Same.dev/Tools.json mandates relative paths rooted at the workspace directory. This design prevents path traversal attacks (such as ../../etc/passwd) by rejecting absolute paths or attempts to navigate above the project root. The tool strictly operates within the sandboxed workspace boundary.

Can an AI assistant delete multiple files at once?

While some implementations like Trae/Builder Tools.json allow deleting multiple files in a single tool call, the repository emphasizes single-file granularity or explicit enumeration of each target. This prevents bulk accidental deletions. Additionally, each file deletion typically requires individual confirmation through the workflow defined in v0 Prompts and Tools/Prompt.txt, ensuring users approve every target specifically.

What happens if a user accidentally approves a file deletion?

The repository implements recovery mechanisms through sandbox environments like the E2B sandbox referenced in Orchids.app/System Prompt.txt. The runtime records a diff of the deleted file's content, allowing the assistant to restore the file using a subsequent create_file call if the user changes their mind. This provides an undo capability even after the deletion operation completes.

Why are shell commands like rm -rf forbidden for file deletion?

Augment Code/claude-4-sonnet-tools.json explicitly forbids shell-based deletions because the dedicated delete_file tool provides tracked, sandboxed, and reversible operations. Shell commands bypass schema-level safety checks (like path validation and graceful failure), escape the runtime's monitoring capabilities, and prevent recovery mechanisms from functioning. The tool-only policy ensures every deletion is subject to the repository's defense-in-depth safety layers.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →