How to Speed Up Video Rendering on Multi-Core Processors in Video-Use
Replace the sequential for loop in extract_all_segments with a ThreadPoolExecutor, add -threads 0 to FFmpeg commands, and utilize hardware-accelerated encoders to distribute video processing across all available CPU cores.
The browser-use/video-use repository renders final videos by extracting segments via FFmpeg, concatenating them losslessly, and applying compositing overlays. By default, the extract_all_segments function in helpers/render.py processes segments sequentially, creating a bottleneck that leaves most CPU cores idle on modern hardware. Implementing parallel extraction strategies and explicit threading controls allows you to speed up video rendering on multi-core processors while maintaining output quality.
Parallelize Segment Extraction with ThreadPoolExecutor
The primary bottleneck resides in extract_all_segments (helpers/render.py#L37-L63), where a standard for loop invokes extract_segment sequentially for each clip range. Since each FFmpeg extraction is independent and I/O-bound, this sequential execution wastes available parallelism on multi-core systems.
Migrate this logic to use concurrent.futures.ThreadPoolExecutor following the proven pattern in helpers/transcribe_batch.py#L85-L98. This distributes segment extraction across all CPU cores, saturating the processor during the extraction phase.
# helpers/render.py – replace the sequential loop in extract_all_segments
from concurrent.futures import ThreadPoolExecutor, as_completed
import os
from pathlib import Path
def extract_all_segments(
edl: dict,
edit_dir: Path,
preview: bool,
draft: bool = False,
) -> list[Path]:
seg_paths: list[Path] = []
clips_dir = edit_dir / "clips"
clips_dir.mkdir(exist_ok=True)
ranges = edl.get("ranges", [])
sources = edl.get("sources", {})
with ThreadPoolExecutor(max_workers=os.cpu_count()) as pool:
future_to_idx = {}
for i, r in enumerate(ranges):
src_name = r["source"]
src_path = Path(sources[src_name]) # resolve_path logic applies here
start = float(r["start"])
end = float(r["end"])
duration = end - start
out_path = clips_dir / f"seg_{i:02d}_{src_name}.mp4"
# Grade filter resolution omitted for brevity
future = pool.submit(
extract_segment,
src_path,
start,
duration,
None, # resolved grade_filter
out_path,
preview=preview,
draft=draft,
)
future_to_idx[future] = (i, out_path)
for fut in as_completed(future_to_idx):
i, out_path = future_to_idx[fut]
fut.result() # propagate exceptions
seg_paths.append(out_path)
print(f" [{i:02d}] extracted → {out_path.name}")
return seg_paths
Leverage FFmpeg's Internal Threading
Even with parallel extraction, individual FFmpeg processes may default to limited internal threading. Force full CPU utilization by appending the -threads 0 flag to enable automatic thread detection based on CPU count.
Modify the command construction in extract_segment (helpers/render.py#L99-L106):
# Inside extract_segment (helpers/render.py)
cmd = [
"ffmpeg", "-y",
"-threads", "0", # Enable auto-threading across CPU cores
"-ss", f"{seg_start:.3f}",
"-i", str(source),
"-t", f"{duration:.3f}",
"-vf", vf,
"-af", af,
"-c:v", "libx264", "-preset", preset, "-crf", crf,
"-pix_fmt", "yuv420p", "-r", "24",
"-c:a", "aac", "-b:a", "192k", "-ar", "48000",
"-movflags", "+faststart",
str(out_path),
]
Apply the same "-threads", "0" addition to the final compositing command in build_final_composite (helpers/render.py#L55-L62).
Optimize Encoding Presets and Hardware Acceleration
The codebase already implements draft and preview presets for rapid iteration (helpers/render.py#L91-L99). Use these faster presets during editing, reserving the slower fast preset only for final delivery to minimize CPU cycles during quality checks.
For hardware acceleration, replace libx264 with GPU-accelerated encoders to offload macro-block calculations from the CPU:
- NVIDIA GPUs: Use
h264_nvenc - macOS: Use
h264_videotoolbox(works on both Apple Silicon and Intel)
Change the -c:v argument in both extract_segment and the final compositing command. This requires an FFmpeg build compiled with the appropriate codecs but frees CPU cores for other parallel tasks.
Preserve Zero-Copy Concatenation
The pipeline already optimizes intermediary processing by concatenating segments with -c copy (helpers/render.py#L66-L71), which streams segments together without re-encoding. Maintain this zero-copy approach for any custom modifications to avoid CPU-intensive re-rendering.
For overlay processing, the build_final_composite function (helpers/render.py#L15-L38) constructs a single filter_complex graph that processes all overlays in one pass. This batch approach minimizes decode cycles and prevents the overhead of multiple intermediate renders.
Summary
To speed up video rendering on multi-core processors in video-use:
- Replace sequential extraction with
ThreadPoolExecutorinextract_all_segmentsto run multiple FFmpeg instances concurrently across available cores - Add
-threads 0to all FFmpeg commands to enable internal multi-threading within each encoding process - Use hardware encoders (
h264_nvenc,h264_videotoolbox) to offload encoding workloads from CPU to GPU - Leverage existing optimizations including lossless
-c copyconcatenation and batch overlay processing viafilter_complex
Frequently Asked Questions
Will parallel segment extraction preserve the correct video order?
Yes. The ThreadPoolExecutor implementation stores segment indices in the future_to_idx mapping, allowing the code to assemble the seg_paths list in the original sequence after all extractions complete. The preserved index i from the EDL range list ensures proper ordering regardless of which segment finishes first.
Does adding -threads 0 significantly increase memory usage?
FFmpeg's internal threading consumes additional memory per encoder instance, but the increase is typically modest compared to the performance gains on multi-core processors. For systems with severe memory constraints, specify a fixed thread count (e.g., -threads 4) instead of 0 to cap resource usage while still benefiting from parallel processing.
Can I use ProcessPoolExecutor instead of ThreadPoolExecutor for segment extraction?
While ProcessPoolExecutor is possible, ThreadPoolExecutor is optimal here because extract_segment launches external FFmpeg subprocesses that release the Python GIL. Threads provide lower overhead than processes for this I/O-bound workload, as demonstrated in the transcribe_batch.py implementation (helpers/transcribe_batch.py#L85-L98).
How do I verify that multi-core rendering is working?
Monitor CPU utilization during rendering using htop (Linux) or Activity Monitor (macOS). With parallel extraction enabled, you should observe multiple FFmpeg processes consuming CPU simultaneously, with total utilization approaching 100% across all logical cores during the extraction phase.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →