Security Considerations When Processing Untrusted Video Inputs in Python

Processing untrusted video inputs requires strict path validation, subprocess isolation, resource limits, and sandboxing of FFmpeg operations to prevent path traversal, command injection, and denial-of-service attacks.

The video-use repository by browser-use demonstrates practical security patterns for handling external video files through helper scripts like transcribe.py and timeline_view.py. When accepting video files from unknown sources, every stage—from initial path resolution to final JSON output—introduces potential attack vectors that must be hardened against malicious input.

Path Traversal Prevention with Path Resolution

All video processing helpers in the repository call Path.resolve() on user-supplied paths before any filesystem operations.

In helpers/transcribe.py (lines 58-60), the code immediately resolves the input:

video = args.video.resolve()

The same pattern appears in helpers/timeline_view.py (lines 361-363):

video = args.video.resolve()

Why this matters: Resolving to an absolute path prevents directory traversal attacks using ../ sequences that could force the script to read or overwrite sensitive files outside the intended directory. However, resolution alone is not sufficient—you should also enforce a whitelist of permissible directories:

def safe_video_path(user_input: str) -> Path:
    base = Path("/trusted/video_inputs").resolve()
    p = Path(user_input).resolve()
    if not str(p).startswith(str(base)):
        raise ValueError("Video path outside allowed directory")
    return p

Command Injection Mitigation via Subprocess Lists

Both transcribe.py and timeline_view.py invoke FFmpeg using a list of arguments rather than a shell string, as seen in transcribe.py (lines 51-55):

cmd = ["ffmpeg", "-y", "-i", str(video_path), ...]
subprocess.run(cmd, check=True, stdout=subprocess.DEVNULL,
               stderr=subprocess.DEVNULL)

timeline_view.py uses the same pattern when extracting frames (lines 50-60). This construction ensures that malicious filenames containing shell metacharacters like ;, &&, or backticks are treated as literal arguments rather than executable commands.

Never use shell=True when processing untrusted paths, as this would reintroduce injection vectors even with seemingly safe input.

Temporary File Isolation and Cleanup

The repository uses tempfile.TemporaryDirectory() for short-lived processing artifacts. In transcribe.py, the transcribe_one function (lines 14-18) creates isolated temporary directories that are automatically cleaned up after processing. Similarly, timeline_view.py creates named temporary WAV files in compute_envelope (lines 74-78).

Temporal isolation prevents leftover artifacts from being reused by attackers and limits the window for race-condition exploits. Always ensure temporary directories are created with restrictive permissions and deleted immediately after use.

Resource Consumption Controls

Untrusted inputs can trigger denial-of-service (DoS) attacks by exhausting CPU, memory, or disk space. The repository addresses this in helpers/transcribe_batch.py (lines 68-71) by enforcing a maximum number of parallel workers via the --workers argument and reporting the total video count before processing.

Additional hardening should include explicit file-size caps:

MAX_WAV_MB = 10
if audio.stat().st_size > MAX_WAV_MB * 1024 * 1024:
    raise RuntimeError("Audio extraction exceeds safe upload size")

FFmpeg Sandboxing and Decoder Hardening

Because FFmpeg parses complex container formats, a crafted video could exploit decoder vulnerabilities. The repository relies heavily on FFmpeg for audio/video extraction in transcribe.py, timeline_view.py, and render.py.

Recommended mitigations:

  • Run FFmpeg inside a sandbox (Docker, firejail, or restricted user namespace)
  • Pin the FFmpeg binary to a known, patched version
  • Disable dangerous codecs via -f (format) or -codec options when feasible

Docker example for sandboxing:

docker run --rm -v "$PWD:/work" \
    --user "$(id -u):$(id -g)" \
    --security-opt=no-new-privileges \
    jrottenberg/ffmpeg:5.1 \
    -i /work/input.mp4 -vf "scale=320:-2" /work/output.jpg

API Key Protection and Network Exposure

transcribe.py uploads extracted audio to the ElevenLabs Scribe API via the call_scribe function (lines 58-62). The API key is loaded from environment variables (lines 33-46), and FFmpeg output is suppressed to prevent credential leakage in logs.

Security considerations:

  • Validate extracted audio size before upload to prevent unexpected billing
  • Ensure API keys never appear in stdout/stderr (already implemented via stdout=subprocess.DEVNULL)
  • Consider using a secrets manager rather than .env files in production environments

Output Sanitization for Downstream Consumers

Transcripts are written as JSON in transcribe.py (lines 23-24):

out_path.write_text(json.dumps(payload, indent=2))

Since this payload originates from the external Scribe service, it should be treated as untrusted. Later consumers parsing this JSON must treat all fields as opaque strings. If transcripts are rendered in web browsers, implement HTML/JS escaping to prevent XSS attacks.

Summary

  • Path.resolve() prevents directory traversal but should be combined with directory whitelisting for defense in depth
  • Subprocess lists (not shell strings) eliminate command injection risks when calling FFmpeg
  • Temporary directories provide temporal isolation and automatic cleanup for processing artifacts
  • Worker limits and size checks mitigate denial-of-service attacks from massive or numerous video files
  • FFmpeg sandboxing is essential due to the attack surface of media decoders
  • API key hygiene requires suppressing logs and validating upload sizes before external service calls
  • JSON output from transcription services must be sanitized before UI rendering to prevent XSS

Frequently Asked Questions

How does video-use prevent command injection when processing filenames?

The repository passes arguments to subprocess.run() as a list rather than a string, which disables shell interpretation. For example, cmd = ["ffmpeg", "-i", str(video_path)] ensures that special characters in filenames are treated as literals rather than shell operators.

Can path traversal attacks still occur despite using Path.resolve()?

While Path.resolve() converts relative paths to absolute paths (neutralizing ../ sequences), it does not enforce directory boundaries. An attacker could still specify /etc/passwd directly. Additional validation ensuring the resolved path starts with a trusted base directory is required for complete protection.

Why is FFmpeg considered a security risk when processing untrusted videos?

FFmpeg contains parsers for hundreds of media formats and codecs, some of which have historically contained memory corruption vulnerabilities. A maliciously crafted video file could exploit these bugs to execute arbitrary code. Running FFmpeg in a sandboxed environment (Docker, restricted user namespace) limits the impact of potential exploits.

How should API keys be protected when uploading processed audio to external services?

Load API keys from environment variables or secrets managers rather than hardcoding them. Suppress command output using stdout=subprocess.DEVNULL and stderr=subprocess.DEVNULL to prevent keys from appearing in logs. Additionally, validate file sizes before upload to prevent quota exhaustion or unexpected billing charges.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →