# How to Process Video Files, Extract Frames, and Set FPS for LingBot-Map

> Learn to process video files for LingBot-Map. Extract frames and set FPS using OpenCV or FFmpeg for 3D reconstruction. Get started now.

- Repository: [Robbyant/lingbot-map](https://github.com/Robbyant/lingbot-map)
- Tags: how-to-guide
- Published: 2026-07-30

---

**LingBot-Map automatically extracts frames from video files using OpenCV or FFmpeg, resamples them to your specified frame rate via the `--fps` argument, and stores them in a sibling folder for 3D reconstruction processing.**

LingBot-Map supports both image directories and raw video files as input sources. When you provide a video path, the repository handles frame extraction, temporal downsampling, and preprocessing automatically. This guide explains how to process video files, extract frames, and set FPS for LingBot-Map using the built-in extraction pipelines.

## Video Input Handling in LingBot-Map

LingBot-Map accepts either a directory of pre-extracted images or a direct video file path. When a video is supplied, the `load_images` function orchestrates the extraction process, converting the video stream into a discrete frame sequence suitable for the reconstruction pipeline.

## OpenCV Frame Extraction (Primary Method)

The default extraction logic resides in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (lines 65-90) within the `load_images` function. This implementation uses OpenCV (`cv2.VideoCapture`) to read video files and process frame sampling according to the repository's internal logic.

### FPS Calculation and Frame Sampling

The extraction interval is calculated based on the source video's native FPS (`src_fps`) and your target FPS:

```python
interval = max(1, round(src_fps / fps))

```

For example, if your source video runs at 30 FPS and you request 10 FPS, the interval becomes 3, preserving every third frame. The function reads the source frame rate directly from the video metadata and applies this downsampling logic to maintain your specified temporal resolution without duplicating frames.

### Frame Storage and Naming Convention

Extracted frames are written as JPEG images to a folder named `<video_name>_frames` created as a sibling directory to the original video file. This location serves as the cache for subsequent preprocessing steps, including the `load_and_preprocess_images` utility which handles cropping and resizing to the configured image size (default 518px).

## FFmpeg-Based Extraction (Benchmark Pipeline)

For batch processing or benchmark datasets, LingBot-Map provides an alternative extractor in [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py) (lines 232-233). This implementation invokes **FFmpeg** with the `-vf fps=<desired_fps>` video filter to extract frames at the specified rate. Like the OpenCV method, it caches frames in a sibling directory and reuses existing extractions on subsequent runs to avoid redundant processing.

## Configuring the Target FPS

You control the extraction frame rate through two primary mechanisms:

- **Command-line interface**: Pass `--fps` when running [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py) (default is 10 FPS)
- **Dataset configuration**: Set the `video_fps` field in your dataset YAML file when using the benchmark loader

The source code automatically computes the sampling interval and manages the extraction process regardless of which method you choose.

## Practical Code Examples

Extract frames at the default 10 FPS and run reconstruction:

```bash
python demo.py \
  --model_path /path/to/checkpoint.pt \
  --video_path /data/my_video.mp4 \
  --fps 10

```

Increase temporal resolution to 20 FPS:

```bash
python demo.py \
  --model_path /path/to/checkpoint.pt \
  --video_path /data/my_video.mp4 \
  --fps 20

```

Use the benchmark loader with FFmpeg at 5 FPS:

```bash
python -m benchmark.run \
  --data_root /data \
  --video_file my_video.mp4 \
  --video_fps 5

```

Programmatic extraction in Python notebooks:

```python
from lingbot_map.demo import load_images

imgs, paths, folder = load_images(
    video_path="my_video.mp4",
    fps=15,
    image_size=518,
    patch_size=14,
)
print(f"Extracted {len(paths)} frames into {folder}")

```

## Summary

- LingBot-Map accepts video files directly and handles frame extraction automatically via `load_images` in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py)
- The OpenCV extractor calculates sampling intervals using `max(1, round(src_fps / fps))` to achieve your target frame rate
- Frames are cached in `<video_name>_frames` directories next to source videos for reuse across runs
- An alternative FFmpeg pipeline exists in [`benchmark/datasets/general.py`](https://github.com/Robbyant/lingbot-map/blob/main/benchmark/datasets/general.py) for batch processing with the `-vf fps=` filter
- Control extraction speed using the `--fps` CLI argument or `video_fps` configuration fields

## Frequently Asked Questions

### What video formats does LingBot-Map support?

LingBot-Map relies on OpenCV's `cv2.VideoCapture` for the primary extraction pipeline, which supports standard formats including MP4, AVI, and MOV. The FFmpeg-based benchmark loader supports any format compatible with your system FFmpeg installation.

### How does LingBot-Map handle target FPS values that don't divide evenly into the source FPS?

The code uses Python's `round()` function within `interval = max(1, round(src_fps / fps))` to determine the nearest integer frame skip. If the ratio isn't exact, the actual output FPS will approximate your target, ensuring you never exceed the requested rate while maintaining consistent temporal spacing.

### Can I reuse extracted frames across multiple reconstruction runs?

Yes. Both extraction methods cache frames in sibling folders (`video_name_frames`). The pipeline checks for existing extractions before processing, allowing you to reuse frames and skip redundant I/O operations on subsequent runs with different model parameters.

### Where does LingBot-Map store the extracted frames?

According to the source code in [`demo.py`](https://github.com/Robbyant/lingbot-map/blob/main/demo.py), frames are written to a directory named `<video_name>_frames` located in the same parent directory as the source video. The function returns this path for reference in downstream preprocessing steps.