# How Video-Use Handles Audio Processing During Video Rendering

> Discover how video-use handles audio processing during video rendering with its four stage pipeline: extraction, concatenation, loudness normalization, and compositing. Learn more!

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: internals
- Published: 2026-06-30

---

**The `video-use` rendering pipeline processes audio through four distinct stages: per-segment extraction with 30ms fade effects, lossless concatenation, optional two-pass loudness normalization to -14 LUFS, and final compositing with stream copying, all orchestrated in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py).**

The `browser-use/video-use` repository provides automated video generation capabilities for browser automation workflows. Understanding its **audio processing during video rendering** requires examining [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py), where FFmpeg commands manage everything from fade-in effects to broadcast-standard loudness normalization.

## The Four-Stage Audio Pipeline

### Stage 1: Per-Segment Extraction with Fade Effects

The `extract_segment()` function in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) handles the initial **audio processing during video rendering** by preventing audible pops and clicks. At lines 88-90, the script constructs an FFmpeg audio filter chain that applies a 30ms fade-in at the start and a 30ms fade-out at the end of each clip:

```python
af = f"afade=t=in:st=0:d=0.03,afade=t=out:st={fade_out_start:.3f}:d=0.03"

```

This filter string gets passed to FFmpeg alongside video encoding parameters. At lines 100-107, the function forces audio re-encoding to AAC at 192 kbit/s with a 48 kHz sample rate, ensuring compatibility while the video stream is encoded with libx264. This stage guarantees that every individual segment has smooth audio transitions before concatenation begins.

### Stage 2: Lossless Audio Concatenation

After individual segments are processed, the `concat_segments()` function (lines 74-78) joins them using FFmpeg's concat demuxer with the `-c copy` flag. This **lossless concatenation** copies both video and audio streams without re-encoding, preserving the fade effects applied in Stage 1 while avoiding generational quality loss. The audio streams remain untouched as they are stitched together in their original AAC format.

### Stage 3: Optional Two-Pass Loudness Normalization

For social media compliance, the pipeline implements an optional two-pass loudness normalization workflow controlled by the `--no-loudnorm` flag. When enabled, `measure_loudness()` (lines 98-104) executes Pass 1 by running FFmpeg with the `loudnorm` filter in measurement mode, parsing the JSON output to capture integrated loudness, true-peak, and loudness range values.

The `apply_loudnorm_two_pass()` function (lines 31-90) then executes Pass 2, applying the measured values to reach the target specification of **-14 LUFS integrated loudness, -1 dBTP true-peak, and LRA 11**—the standard for YouTube, TikTok, and Instagram. In draft or preview mode, the pipeline skips the measurement pass and uses a single-pass approximation for faster rendering.

### Stage 4: Final Audio Compositing

The `build_final_composite()` function (lines 59-64) handles the final stage by mapping the audio stream from the base concatenated clip using `-map 0:a` and copying it directly with `-c:a copy`. If loudness normalization was applied in Stage 3, the audio stream is already normalized; otherwise, the original concatenated audio passes through unchanged. No further audio processing occurs during overlay compositing.

## FFmpeg Implementation Details

The following Python snippets demonstrate the core FFmpeg command construction for **audio processing during video rendering**:

**Per-segment extraction with fade effects:**

```python
af = f"afade=t=in:st=0:d=0.03,afade=t=out:st={fade_out_start:.3f}:d=0.03"
cmd = [
    "ffmpeg", "-y",
    "-ss", f"{seg_start:.3f}",
    "-i", str(source),
    "-t", f"{duration:.3f}",
    "-vf", vf,
    "-af", af,
    "-c:v", "libx264", "-preset", preset, "-crf", crf,
    "-c:a", "aac", "-b:a", "192k", "-ar", "48000",
    "-movflags", "+faststart",
    str(out_path),
]

```

**Two-pass loudness normalization workflow:**

```python

# Pass 1 – measurement

measurement = measure_loudness(input_path)

# Pass 2 – apply measured values

filter_str = (
    f"loudnorm=I={LOUDNORM_I}:TP={LOUDNORM_TP}:LRA={LOUDNORM_LRA}"
    f":measured_I={measurement['input_i']}"
    f":measured_TP={measurement['input_tp']}"
    f":measured_LRA={measurement['input_lra']}"
    f":measured_thresh={measurement['input_thresh']}"
    f":offset={measurement['target_offset']}"
    f":linear=true"
)
cmd = [
    "ffmpeg", "-y", "-hide_banner", "-nostats",
    "-i", str(input_path),
    "-c:v", "copy",
    "-af", filter_str,
    "-c:a", "aac", "-b:a", "192k", "-ar", "48000",
    "-movflags", "+faststart",
    str(output_path),
]

```

**Final compositing with audio stream copying:**

```python
cmd = [
    "ffmpeg", "-y",
    *inputs,                     # base + any overlay inputs

    "-filter_complex", filter_complex,
    "-map", out_label,          # video

    "-map", "0:a",              # audio from base

    "-c:v", "libx264", "-preset", "fast", "-crf", "18",
    "-pix_fmt", "yuv420p",
    "-c:a", "copy",             # keep audio unchanged

    "-movflags", "+faststart",
    str(out_path),
]

```

## Command-Line Audio Controls

The rendering behavior can be modified through command-line flags defined in the repository's entry points:

- **`--no-loudnorm`**: Skips the two-pass loudness normalization entirely, preserving the original audio levels from the concatenated segments.
- **`--draft`** or **`--preview`**: Enables single-pass loudness approximation for faster rendering during editing iterations, bypassing the measurement phase.

## Summary

- The **four-stage pipeline** in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) handles audio through extraction, concatenation, optional normalization, and final compositing.
- **30ms fade effects** applied in `extract_segment()` eliminate audio pops at segment boundaries.
- **Lossless concatenation** via `concat_segments()` preserves audio quality using `-c copy`.
- **Two-pass loudnorm** targets -14 LUFS/-1 dBTP for social media compliance when normalization is enabled.
- Final output uses **AAC 192 kbit/s at 48 kHz** regardless of processing path.

## Frequently Asked Questions

### Why does Video-Use apply 30ms fade effects to audio segments?

The 30ms fade-in and fade-out effects prevent audible pops and clicks that occur when audio waveforms start or end abruptly at non-zero crossings. This is implemented in `extract_segment()` at lines 88-90 using FFmpeg's `afade` filter, ensuring smooth transitions between clips during the concatenation stage.

### What is the purpose of two-pass loudness normalization?

Two-pass normalization first measures the audio's current loudness characteristics (integrated loudness, true-peak, and loudness range) and then applies precise corrections to meet the -14 LUFS target. This approach ensures broadcast-compliant audio levels for social platforms, implemented in `measure_loudness()` and `apply_loudnorm_two_pass()` in [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py).

### How can I disable loudness normalization during rendering?

Pass the `--no-loudnorm` flag when running the render command. This skips both the measurement and correction passes, allowing the audio to remain at its original levels from the concatenation stage while still maintaining the 30ms fades applied during segment extraction.

### What audio codec settings does Video-Use use for final output?

The pipeline encodes audio to **AAC (Advanced Audio Coding)** at **192 kbit/s** with a **48 kHz** sample rate. These settings are hardcoded in `extract_segment()` (lines 100-107) and reused during loudness normalization, ensuring consistent compatibility across all major video platforms.