How Video-Use Handles Audio Processing During Video Rendering

The video-use rendering pipeline processes audio through four distinct stages: per-segment extraction with 30ms fade effects, lossless concatenation, optional two-pass loudness normalization to -14 LUFS, and final compositing with stream copying, all orchestrated in helpers/render.py.

The browser-use/video-use repository provides automated video generation capabilities for browser automation workflows. Understanding its audio processing during video rendering requires examining helpers/render.py, where FFmpeg commands manage everything from fade-in effects to broadcast-standard loudness normalization.

The Four-Stage Audio Pipeline

Stage 1: Per-Segment Extraction with Fade Effects

The extract_segment() function in helpers/render.py handles the initial audio processing during video rendering by preventing audible pops and clicks. At lines 88-90, the script constructs an FFmpeg audio filter chain that applies a 30ms fade-in at the start and a 30ms fade-out at the end of each clip:

af = f"afade=t=in:st=0:d=0.03,afade=t=out:st={fade_out_start:.3f}:d=0.03"

This filter string gets passed to FFmpeg alongside video encoding parameters. At lines 100-107, the function forces audio re-encoding to AAC at 192 kbit/s with a 48 kHz sample rate, ensuring compatibility while the video stream is encoded with libx264. This stage guarantees that every individual segment has smooth audio transitions before concatenation begins.

Stage 2: Lossless Audio Concatenation

After individual segments are processed, the concat_segments() function (lines 74-78) joins them using FFmpeg's concat demuxer with the -c copy flag. This lossless concatenation copies both video and audio streams without re-encoding, preserving the fade effects applied in Stage 1 while avoiding generational quality loss. The audio streams remain untouched as they are stitched together in their original AAC format.

Stage 3: Optional Two-Pass Loudness Normalization

For social media compliance, the pipeline implements an optional two-pass loudness normalization workflow controlled by the --no-loudnorm flag. When enabled, measure_loudness() (lines 98-104) executes Pass 1 by running FFmpeg with the loudnorm filter in measurement mode, parsing the JSON output to capture integrated loudness, true-peak, and loudness range values.

The apply_loudnorm_two_pass() function (lines 31-90) then executes Pass 2, applying the measured values to reach the target specification of -14 LUFS integrated loudness, -1 dBTP true-peak, and LRA 11—the standard for YouTube, TikTok, and Instagram. In draft or preview mode, the pipeline skips the measurement pass and uses a single-pass approximation for faster rendering.

Stage 4: Final Audio Compositing

The build_final_composite() function (lines 59-64) handles the final stage by mapping the audio stream from the base concatenated clip using -map 0:a and copying it directly with -c:a copy. If loudness normalization was applied in Stage 3, the audio stream is already normalized; otherwise, the original concatenated audio passes through unchanged. No further audio processing occurs during overlay compositing.

FFmpeg Implementation Details

The following Python snippets demonstrate the core FFmpeg command construction for audio processing during video rendering:

Per-segment extraction with fade effects:

af = f"afade=t=in:st=0:d=0.03,afade=t=out:st={fade_out_start:.3f}:d=0.03"
cmd = [
    "ffmpeg", "-y",
    "-ss", f"{seg_start:.3f}",
    "-i", str(source),
    "-t", f"{duration:.3f}",
    "-vf", vf,
    "-af", af,
    "-c:v", "libx264", "-preset", preset, "-crf", crf,
    "-c:a", "aac", "-b:a", "192k", "-ar", "48000",
    "-movflags", "+faststart",
    str(out_path),
]

Two-pass loudness normalization workflow:


# Pass 1 – measurement

measurement = measure_loudness(input_path)

# Pass 2 – apply measured values

filter_str = (
    f"loudnorm=I={LOUDNORM_I}:TP={LOUDNORM_TP}:LRA={LOUDNORM_LRA}"
    f":measured_I={measurement['input_i']}"
    f":measured_TP={measurement['input_tp']}"
    f":measured_LRA={measurement['input_lra']}"
    f":measured_thresh={measurement['input_thresh']}"
    f":offset={measurement['target_offset']}"
    f":linear=true"
)
cmd = [
    "ffmpeg", "-y", "-hide_banner", "-nostats",
    "-i", str(input_path),
    "-c:v", "copy",
    "-af", filter_str,
    "-c:a", "aac", "-b:a", "192k", "-ar", "48000",
    "-movflags", "+faststart",
    str(output_path),
]

Final compositing with audio stream copying:

cmd = [
    "ffmpeg", "-y",
    *inputs,                     # base + any overlay inputs

    "-filter_complex", filter_complex,
    "-map", out_label,          # video

    "-map", "0:a",              # audio from base

    "-c:v", "libx264", "-preset", "fast", "-crf", "18",
    "-pix_fmt", "yuv420p",
    "-c:a", "copy",             # keep audio unchanged

    "-movflags", "+faststart",
    str(out_path),
]

Command-Line Audio Controls

The rendering behavior can be modified through command-line flags defined in the repository's entry points:

  • --no-loudnorm: Skips the two-pass loudness normalization entirely, preserving the original audio levels from the concatenated segments.
  • --draft or --preview: Enables single-pass loudness approximation for faster rendering during editing iterations, bypassing the measurement phase.

Summary

  • The four-stage pipeline in helpers/render.py handles audio through extraction, concatenation, optional normalization, and final compositing.
  • 30ms fade effects applied in extract_segment() eliminate audio pops at segment boundaries.
  • Lossless concatenation via concat_segments() preserves audio quality using -c copy.
  • Two-pass loudnorm targets -14 LUFS/-1 dBTP for social media compliance when normalization is enabled.
  • Final output uses AAC 192 kbit/s at 48 kHz regardless of processing path.

Frequently Asked Questions

Why does Video-Use apply 30ms fade effects to audio segments?

The 30ms fade-in and fade-out effects prevent audible pops and clicks that occur when audio waveforms start or end abruptly at non-zero crossings. This is implemented in extract_segment() at lines 88-90 using FFmpeg's afade filter, ensuring smooth transitions between clips during the concatenation stage.

What is the purpose of two-pass loudness normalization?

Two-pass normalization first measures the audio's current loudness characteristics (integrated loudness, true-peak, and loudness range) and then applies precise corrections to meet the -14 LUFS target. This approach ensures broadcast-compliant audio levels for social platforms, implemented in measure_loudness() and apply_loudnorm_two_pass() in helpers/render.py.

How can I disable loudness normalization during rendering?

Pass the --no-loudnorm flag when running the render command. This skips both the measurement and correction passes, allowing the audio to remain at its original levels from the concatenation stage while still maintaining the 30ms fades applied during segment extraction.

What audio codec settings does Video-Use use for final output?

The pipeline encodes audio to AAC (Advanced Audio Coding) at 192 kbit/s with a 48 kHz sample rate. These settings are hardcoded in extract_segment() (lines 100-107) and reused during loudness normalization, ensuring consistent compatibility across all major video platforms.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →