How to Process Video Files, Extract Frames, and Set FPS for LingBot-Map

LingBot-Map automatically extracts frames from video files using OpenCV or FFmpeg, resamples them to your specified frame rate via the --fps argument, and stores them in a sibling folder for 3D reconstruction processing.

LingBot-Map supports both image directories and raw video files as input sources. When you provide a video path, the repository handles frame extraction, temporal downsampling, and preprocessing automatically. This guide explains how to process video files, extract frames, and set FPS for LingBot-Map using the built-in extraction pipelines.

Video Input Handling in LingBot-Map

LingBot-Map accepts either a directory of pre-extracted images or a direct video file path. When a video is supplied, the load_images function orchestrates the extraction process, converting the video stream into a discrete frame sequence suitable for the reconstruction pipeline.

OpenCV Frame Extraction (Primary Method)

The default extraction logic resides in demo.py (lines 65-90) within the load_images function. This implementation uses OpenCV (cv2.VideoCapture) to read video files and process frame sampling according to the repository's internal logic.

FPS Calculation and Frame Sampling

The extraction interval is calculated based on the source video's native FPS (src_fps) and your target FPS:

interval = max(1, round(src_fps / fps))

For example, if your source video runs at 30 FPS and you request 10 FPS, the interval becomes 3, preserving every third frame. The function reads the source frame rate directly from the video metadata and applies this downsampling logic to maintain your specified temporal resolution without duplicating frames.

Frame Storage and Naming Convention

Extracted frames are written as JPEG images to a folder named <video_name>_frames created as a sibling directory to the original video file. This location serves as the cache for subsequent preprocessing steps, including the load_and_preprocess_images utility which handles cropping and resizing to the configured image size (default 518px).

FFmpeg-Based Extraction (Benchmark Pipeline)

For batch processing or benchmark datasets, LingBot-Map provides an alternative extractor in benchmark/datasets/general.py (lines 232-233). This implementation invokes FFmpeg with the -vf fps=<desired_fps> video filter to extract frames at the specified rate. Like the OpenCV method, it caches frames in a sibling directory and reuses existing extractions on subsequent runs to avoid redundant processing.

Configuring the Target FPS

You control the extraction frame rate through two primary mechanisms:

  • Command-line interface: Pass --fps when running demo.py (default is 10 FPS)
  • Dataset configuration: Set the video_fps field in your dataset YAML file when using the benchmark loader

The source code automatically computes the sampling interval and manages the extraction process regardless of which method you choose.

Practical Code Examples

Extract frames at the default 10 FPS and run reconstruction:

python demo.py \
  --model_path /path/to/checkpoint.pt \
  --video_path /data/my_video.mp4 \
  --fps 10

Increase temporal resolution to 20 FPS:

python demo.py \
  --model_path /path/to/checkpoint.pt \
  --video_path /data/my_video.mp4 \
  --fps 20

Use the benchmark loader with FFmpeg at 5 FPS:

python -m benchmark.run \
  --data_root /data \
  --video_file my_video.mp4 \
  --video_fps 5

Programmatic extraction in Python notebooks:

from lingbot_map.demo import load_images

imgs, paths, folder = load_images(
    video_path="my_video.mp4",
    fps=15,
    image_size=518,
    patch_size=14,
)
print(f"Extracted {len(paths)} frames into {folder}")

Summary

  • LingBot-Map accepts video files directly and handles frame extraction automatically via load_images in demo.py
  • The OpenCV extractor calculates sampling intervals using max(1, round(src_fps / fps)) to achieve your target frame rate
  • Frames are cached in <video_name>_frames directories next to source videos for reuse across runs
  • An alternative FFmpeg pipeline exists in benchmark/datasets/general.py for batch processing with the -vf fps= filter
  • Control extraction speed using the --fps CLI argument or video_fps configuration fields

Frequently Asked Questions

What video formats does LingBot-Map support?

LingBot-Map relies on OpenCV's cv2.VideoCapture for the primary extraction pipeline, which supports standard formats including MP4, AVI, and MOV. The FFmpeg-based benchmark loader supports any format compatible with your system FFmpeg installation.

How does LingBot-Map handle target FPS values that don't divide evenly into the source FPS?

The code uses Python's round() function within interval = max(1, round(src_fps / fps)) to determine the nearest integer frame skip. If the ratio isn't exact, the actual output FPS will approximate your target, ensuring you never exceed the requested rate while maintaining consistent temporal spacing.

Can I reuse extracted frames across multiple reconstruction runs?

Yes. Both extraction methods cache frames in sibling folders (video_name_frames). The pipeline checks for existing extractions before processing, allowing you to reuse frames and skip redundant I/O operations on subsequent runs with different model parameters.

Where does LingBot-Map store the extracted frames?

According to the source code in demo.py, frames are written to a directory named <video_name>_frames located in the same parent directory as the source video. The function returns this path for reference in downstream preprocessing steps.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →