# How to Find the Main Entry Point and Core Modules in the video-use Repository

> Discover the main entry point and core modules within the browser-use/video-use repository. Learn how helper scripts initiate the video editing pipeline.

- Repository: [Browser Use/video-use](https://github.com/browser-use/video-use)
- Tags: how-to-guide
- Published: 2026-07-01

---

**The video-use repository does not expose a traditional console script entry point; instead, execution flows through six standalone helper scripts in the `helpers/` directory, each containing an `if __name__ == "__main__":` block that serves as the practical entry point for the video editing pipeline.**

The **video-use** skill is a Python package designed to run inside a Claude-Code (or other LLM-driven) agent for automated video editing. Unlike conventional Python packages that define entry points in [`setup.py`](https://github.com/browser-use/video-use/blob/main/setup.py) or [`pyproject.toml`](https://github.com/browser-use/video-use/blob/main/pyproject.toml), this repository operates as a collection of stage-specific utilities that the host agent invokes sequentially. Understanding where the execution begins requires examining the helper scripts that handle transcription, packing, rendering, and grading.

## Understanding the Repository Architecture

The repository follows a **pipeline-based architecture** rather than a monolithic application structure. According to the `browser-use/video-use` source code, the package is loaded by the host agent via a directory link at `~/.claude/skills/video-use`, which means there is no single `main()` function to import. Instead, the workflow is divided into discrete stages, each implemented by a dedicated helper script that can be executed independently from the command line.

The core execution flow progresses through these stages:

1. **Transcribe** – Extracts audio and generates JSON transcripts using the ElevenLabs Scribe API
2. **Pack Transcripts** – Consolidates transcripts into a single [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md) file for LLM consumption
3. **Reasoning/Planning** – The LLM processes the packed transcripts and generates an Edit-Decision-List (EDL) as specified in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md)
4. **Render** – Applies the EDL to produce the final video using `ffmpeg`
5. **Self-Evaluation** – Generates preview frames for visual verification
6. **Grading** – Evaluates output quality and determines if re-rendering is necessary (max 3 attempts)

## The Pipeline-Based Entry Points

Each stage in the editing pipeline corresponds to a specific file in the `helpers/` directory. These files represent the **core modules** of the repository and contain the executable logic for the video processing workflow.

### Transcribe Stage

The [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py) file serves as the typical starting point for the pipeline. It extracts audio from raw video files, sends the data to the ElevenLabs Scribe API, and caches the resulting JSON transcripts in the `edit/transcripts/` directory.

```python

# From helpers/transcribe.py

if __name__ == "__main__":
    # Handles single video transcription

    # Creates: edit/transcripts/{video_name}.json

```

For batch processing, [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py) extends this functionality to handle multiple videos in a directory.

### Pack Transcripts Stage

Once transcription is complete, [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py) consolidates all individual JSON transcripts into a single markdown file named [`takes_packed.md`](https://github.com/browser-use/video-use/blob/main/takes_packed.md). This file (approximately 12KB) is the primary input that the LLM reads to understand the video content and make editing decisions.

```bash
python helpers/pack_transcripts.py ./edit

```

### Render Stage

The [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py) module constitutes the primary output stage, turning the LLM-generated EDL into the final video file. It executes `ffmpeg` commands to apply cuts, transitions, and optional animation overlays, producing `edit/final.mp4` along with intermediate files in `edit/render/`.

```bash
python helpers/render.py ./edit

```

### Evaluation and Grading

The final stages utilize [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py) for generating diagnostic PNG images of specific time ranges (used during self-evaluation loops) and [`helpers/grade.py`](https://github.com/browser-use/video-use/blob/main/helpers/grade.py) for scoring the rendered output. The grading module determines whether the video meets quality thresholds or requires re-rendering (up to 3 attempts).

```bash

# Generate timeline visualization for verification

python helpers/timeline_view.py ./edit 12.3 15.7

```

## How to Locate Entry Points in the Source Code

To programmatically identify all executable modules in the repository, search for the standard Python entry point pattern:

```bash
rg "if __name__ == \"__main__\"" -g "*.py"

```

This search returns six helper files that serve as the practical **main entry points**:

- [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py)
- [`helpers/transcribe_batch.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe_batch.py)
- [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py)
- [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py)
- [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py)
- [`helpers/grade.py`](https://github.com/browser-use/video-use/blob/main/helpers/grade.py)

Additionally, examining [`pyproject.toml`](https://github.com/browser-use/video-use/blob/main/pyproject.toml) confirms that no console-script entry points are defined, verifying that execution is manual via the helper scripts. The [`README.md`](https://github.com/browser-use/video-use/blob/main/README.md) and [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) files provide the high-level workflow documentation that points developers to these scripts as the canonical entry points.

## Practical Usage Examples

### Transcribing a Single Video

Start the pipeline by processing an individual raw take:

```bash
python helpers/transcribe.py path/to/raw_take.mp4

```

**Result**: Creates [`edit/transcripts/raw_take.json`](https://github.com/browser-use/video-use/blob/main/edit/transcripts/raw_take.json) and caches the extracted audio file.

### Packing Transcripts for LLM Analysis

Consolidate all transcripts before invoking the LLM reasoning stage:

```bash
python helpers/pack_transcripts.py ./edit

```

**Result**: Generates [`edit/takes_packed.md`](https://github.com/browser-use/video-use/blob/main/edit/takes_packed.md) ready for the LLM agent.

### Rendering the Final Video

After the LLM produces an EDL (stored in the edit directory), generate the final output:

```bash
python helpers/render.py ./edit

```

**Result**: Produces `edit/final.mp4` with all edits applied.

### Running a Complete Batch Workflow

Process multiple videos through the entire pipeline:

```bash

# Transcribe all videos in directory

python helpers/transcribe_batch.py ./videos

# Prepare consolidated transcript

python helpers/pack_transcripts.py ./videos/edit

# LLM reasoning happens in the host agent (Claude Code) using SKILL.md

# Render final output

python helpers/render.py ./videos/edit

# Evaluate quality

python helpers/grade.py ./videos/edit

```

## Summary

- The **video-use** repository lacks a traditional single entry point; instead, it distributes functionality across six helper scripts in the `helpers/` directory.
- Each core module contains an `if __name__ == "__main__":` block, making them directly executable from the command line.
- The standard workflow begins with [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py), proceeds through [`helpers/pack_transcripts.py`](https://github.com/browser-use/video-use/blob/main/helpers/pack_transcripts.py), and concludes with [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py), [`helpers/timeline_view.py`](https://github.com/browser-use/video-use/blob/main/helpers/timeline_view.py), and [`helpers/grade.py`](https://github.com/browser-use/video-use/blob/main/helpers/grade.py).
- The [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) file defines the LLM reasoning logic that bridges the transcript packing and rendering stages, but execution is driven by the host agent rather than Python code.
- No console scripts are defined in [`pyproject.toml`](https://github.com/browser-use/video-use/blob/main/pyproject.toml), confirming that the helper scripts are the intended entry points for both developers and automated agents.

## Frequently Asked Questions

### Where is the main function in video-use?

There is no single `main()` function or [`__main__.py`](https://github.com/browser-use/video-use/blob/main/__main__.py) file. According to the `browser-use/video-use` source code, the repository is designed as a skill package loaded by an LLM agent. The **practical entry points** are the `if __name__ == "__main__":` blocks found in [`helpers/transcribe.py`](https://github.com/browser-use/video-use/blob/main/helpers/transcribe.py), [`helpers/render.py`](https://github.com/browser-use/video-use/blob/main/helpers/render.py), and the other helper scripts, which allow each pipeline stage to run as a standalone command.

### How do I run the video-use pipeline from the command line?

Execute the helper scripts sequentially from the repository root. Start with `python helpers/transcribe.py <video_file>`, then run `python helpers/pack_transcripts.py <edit_dir>`. After the LLM generates an EDL (using the logic in [`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md)), run `python helpers/render.py <edit_dir>` to produce the final video. Each script is designed to be invoked independently without importing a central package module.

### What is the purpose of the SKILL.md file?

[`SKILL.md`](https://github.com/browser-use/video-use/blob/main/SKILL.md) serves as the **specification document** for the LLM agent. It contains the 12 production rules and the reasoning logic that guides the agent from the packed transcripts to the Edit-Decision-List (EDL). While it is not executable Python code, it defines the cognitive entry point for the AI-driven portion of the workflow that occurs between the packing and rendering stages.

### Why are there no entry points defined in pyproject.toml?

The [`pyproject.toml`](https://github.com/browser-use/video-use/blob/main/pyproject.toml) file in the video-use repository defines package metadata and dependencies but **omits console-script entry points** because the tool is designed to be loaded as a skill directory by Claude Code or similar agents. The execution model relies on direct script invocation via the `helpers/` modules rather than installed command-line utilities, reflecting its architecture as an agent-plugin rather than a standalone CLI tool.