Per-Segment Extract and Lossless Concatenation Pipeline in video-use
The video-use rendering engine processes footage by first extracting individual clips with color grading and audio fades, then concatenating them losslessly using FFmpeg's concat demuxer to preserve exact quality before final compositing.
The video-use repository implements a three-phase rendering pipeline that separates clip processing from final assembly. At its core, the per-segment extract and lossless concatenation pipeline found in helpers/render.py extracts each edit decision list (EDL) range individually, applies transformations, and stitches them together without re-encoding to maintain pristine quality.
Per-Segment Extraction with extract_segment
The extract_segment function in helpers/render.py (lines 152-190) handles individual clip extraction with precise frame-level control. This function processes each segment defined in the EDL through a series of FFmpeg filter chains to prepare footage for assembly.
Fast Seeking and Input Handling
Accurate seeking is critical for frame-perfect extraction. The pipeline places the -ss flag before the input file to enable FFmpeg's fast seek mechanism, avoiding slow decoding of unwanted frames:
"-ss", f"{seg_start:.3f}", "-i", str(source)
This positioning ensures the decoder jumps directly to the segment start time, significantly accelerating processing for long source files.
HDR Tone Mapping and Portrait Scaling
The pipeline automatically detects HDR content and applies tone mapping when is_hdr_source returns true. The TONEMAP_CHAIN filter is prepended to the video filter graph to convert HDR to SDR while preserving color accuracy.
For mobile footage, the is_portrait_source check determines orientation and applies appropriate scaling:
- Landscape:
scale=1920:-2 - Portrait:
scale=-2:1920
This maintains 1920px on the longest edge while preserving aspect ratios.
Color Grading and Audio Fades
The resolved grade_filter (whether preset, raw FFmpeg filter, or auto-grade output) is appended to the video filter chain via vf_parts.append(grade_filter).
To prevent audio pops at clip boundaries, the pipeline applies 30-millisecond fade curves:
af = f"afade=t=in:st=0:d=0.03,afade=t=out:st={fade_out_start:.3f}:d=0.03"
These fades are applied via the -af flag during the extraction pass.
Quality Presets and Encoding
extract_segment selects from three quality ladders based on the render mode:
- Draft: 720p, ultrafast preset, CRF 28
- Preview: 1080p, medium preset, CRF 22
- Final: 1080p, fast preset, CRF 20
The function encodes to libx264 with yuv420p pixel format, ensuring compatibility with the subsequent lossless concatenation step.
Lossless Concatenation via concat_segments
After all segments are extracted, the concat_segments function (lines 267-283 in helpers/render.py) assembles them without re-encoding. This preservation step is crucial for maintaining the exact quality established during extraction.
Creating the Concatenation List
The function generates a temporary text file named _concat.txt containing absolute paths to each extracted segment:
concat_list.write_text("".join(f"file '{p.resolve()}'\n" for p in segment_paths))
This list follows FFmpeg's concat demuxer format, with one file entry per line.
FFmpeg Copy Mode and Stream Preservation
The concatenation command uses -c copy to stream-copy both video and audio without decoding:
ffmpeg -f concat -safe 0 -i _concat.txt -c copy -movflags +faststart output.mp4
Because all extracted segments share identical codec parameters (libx264, yuv420p, and matching resolution), FFmpeg can stitch them instantaneously while preserving exact timestamps and quality. The -movflags +faststart option ensures web-ready playback by moving the moov atom to the file header.
After successful concatenation, the temporary list file is removed via concat_list.unlink(missing_ok=True).
Orchestrating the Full Pipeline
The extract_all_segments helper (lines 214-260) coordinates the extraction phase by iterating over every range in the EDL. It resolves per-segment grades (including auto-grading when grade: "auto" is specified) and invokes extract_segment for each clip.
From EDL to Final Base Video
In practice, the pipeline executes as follows:
# Extract segments and concatenate (preview quality)
python helpers/render.py my_edl.json -o final.mp4 --preview
# Internal flow:
# 1. extract_all_segments → extract_segment (per-clip processing)
# 2. concat_segments → base_preview.mp4 (lossless assembly)
# 3. build_final_composite (overlays/subtitles)
When run without --preview or --draft flags, the pipeline uses final quality settings (1080p, CRF 20). The concatenated base video is stored as base.mp4 in the EDL directory, ready for the final compositing stage where overlays and subtitles are applied.
Summary
- Per-segment extraction in
helpers/render.pyprocesses each EDL range individually with color grading, HDR tone mapping, and 30ms audio fades. - Fast seeking is achieved by placing
-ssbefore the input file, while portrait orientation triggers height-based scaling to maintain 1920px on the longest edge. - Lossless concatenation uses FFmpeg's concat demuxer with
-c copyto assemble clips without quality degradation, requiring identical codec parameters across all segments. - Quality ladders provide three presets (draft, preview, final) with varying CRF values and resolutions to balance speed against quality during iteration.
- Automatic cleanup removes temporary concat list files after successful assembly, leaving only the final base video for compositing.
Frequently Asked Questions
What is the purpose of the 30ms audio fades in video-use?
The 30-millisecond fade-in and fade-out prevents audible pops and clicks at clip boundaries. These artifacts occur when audio waveforms start or end abruptly at non-zero crossings; the afade filter ramps the amplitude smoothly from and to silence over 30ms, ensuring clean transitions between concatenated segments.
Why does video-use use -c copy instead of re-encoding during concatenation?
The -c copy flag performs stream copying without decoding and re-encoding, which preserves the exact bitstream quality established during the extraction phase. Since all segments are already encoded with identical parameters (libx264, yuv420p, matching resolution), FFmpeg can concatenate them instantly while avoiding generation loss and significantly reducing processing time.
How does the pipeline handle HDR source footage?
When is_hdr_source detects HDR content, the pipeline prepends the TONEMAP_CHAIN filter to the video filter graph before color grading. This converts high dynamic range footage to standard dynamic range using tone mapping algorithms, ensuring consistent color reproduction across the final output without clipping highlights or crushing shadows.
What is the difference between the draft, preview, and final quality presets?
The three presets trade quality for processing speed: Draft (720p, ultrafast, CRF 28) provides quick iteration for rough cuts; Preview (1080p, medium, CRF 22) offers near-final quality for client review; Final (1080p, fast, CRF 20) delivers broadcast-quality output with optimized encoding. Each preset adjusts both resolution and encoding parameters to match the intended use case.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →