How to Find the Main Entry Point and Core Modules in the video-use Repository

The video-use repository does not expose a traditional console script entry point; instead, execution flows through six standalone helper scripts in the helpers/ directory, each containing an if __name__ == "__main__": block that serves as the practical entry point for the video editing pipeline.

The video-use skill is a Python package designed to run inside a Claude-Code (or other LLM-driven) agent for automated video editing. Unlike conventional Python packages that define entry points in setup.py or pyproject.toml, this repository operates as a collection of stage-specific utilities that the host agent invokes sequentially. Understanding where the execution begins requires examining the helper scripts that handle transcription, packing, rendering, and grading.

Understanding the Repository Architecture

The repository follows a pipeline-based architecture rather than a monolithic application structure. According to the browser-use/video-use source code, the package is loaded by the host agent via a directory link at ~/.claude/skills/video-use, which means there is no single main() function to import. Instead, the workflow is divided into discrete stages, each implemented by a dedicated helper script that can be executed independently from the command line.

The core execution flow progresses through these stages:

  1. Transcribe – Extracts audio and generates JSON transcripts using the ElevenLabs Scribe API
  2. Pack Transcripts – Consolidates transcripts into a single takes_packed.md file for LLM consumption
  3. Reasoning/Planning – The LLM processes the packed transcripts and generates an Edit-Decision-List (EDL) as specified in SKILL.md
  4. Render – Applies the EDL to produce the final video using ffmpeg
  5. Self-Evaluation – Generates preview frames for visual verification
  6. Grading – Evaluates output quality and determines if re-rendering is necessary (max 3 attempts)

The Pipeline-Based Entry Points

Each stage in the editing pipeline corresponds to a specific file in the helpers/ directory. These files represent the core modules of the repository and contain the executable logic for the video processing workflow.

Transcribe Stage

The helpers/transcribe.py file serves as the typical starting point for the pipeline. It extracts audio from raw video files, sends the data to the ElevenLabs Scribe API, and caches the resulting JSON transcripts in the edit/transcripts/ directory.


# From helpers/transcribe.py

if __name__ == "__main__":
    # Handles single video transcription

    # Creates: edit/transcripts/{video_name}.json

For batch processing, helpers/transcribe_batch.py extends this functionality to handle multiple videos in a directory.

Pack Transcripts Stage

Once transcription is complete, helpers/pack_transcripts.py consolidates all individual JSON transcripts into a single markdown file named takes_packed.md. This file (approximately 12KB) is the primary input that the LLM reads to understand the video content and make editing decisions.

python helpers/pack_transcripts.py ./edit

Render Stage

The helpers/render.py module constitutes the primary output stage, turning the LLM-generated EDL into the final video file. It executes ffmpeg commands to apply cuts, transitions, and optional animation overlays, producing edit/final.mp4 along with intermediate files in edit/render/.

python helpers/render.py ./edit

Evaluation and Grading

The final stages utilize helpers/timeline_view.py for generating diagnostic PNG images of specific time ranges (used during self-evaluation loops) and helpers/grade.py for scoring the rendered output. The grading module determines whether the video meets quality thresholds or requires re-rendering (up to 3 attempts).


# Generate timeline visualization for verification

python helpers/timeline_view.py ./edit 12.3 15.7

How to Locate Entry Points in the Source Code

To programmatically identify all executable modules in the repository, search for the standard Python entry point pattern:

rg "if __name__ == \"__main__\"" -g "*.py"

This search returns six helper files that serve as the practical main entry points:

Additionally, examining pyproject.toml confirms that no console-script entry points are defined, verifying that execution is manual via the helper scripts. The README.md and SKILL.md files provide the high-level workflow documentation that points developers to these scripts as the canonical entry points.

Practical Usage Examples

Transcribing a Single Video

Start the pipeline by processing an individual raw take:

python helpers/transcribe.py path/to/raw_take.mp4

Result: Creates edit/transcripts/raw_take.json and caches the extracted audio file.

Packing Transcripts for LLM Analysis

Consolidate all transcripts before invoking the LLM reasoning stage:

python helpers/pack_transcripts.py ./edit

Result: Generates edit/takes_packed.md ready for the LLM agent.

Rendering the Final Video

After the LLM produces an EDL (stored in the edit directory), generate the final output:

python helpers/render.py ./edit

Result: Produces edit/final.mp4 with all edits applied.

Running a Complete Batch Workflow

Process multiple videos through the entire pipeline:


# Transcribe all videos in directory

python helpers/transcribe_batch.py ./videos

# Prepare consolidated transcript

python helpers/pack_transcripts.py ./videos/edit

# LLM reasoning happens in the host agent (Claude Code) using SKILL.md

# Render final output

python helpers/render.py ./videos/edit

# Evaluate quality

python helpers/grade.py ./videos/edit

Summary

  • The video-use repository lacks a traditional single entry point; instead, it distributes functionality across six helper scripts in the helpers/ directory.
  • Each core module contains an if __name__ == "__main__": block, making them directly executable from the command line.
  • The standard workflow begins with helpers/transcribe.py, proceeds through helpers/pack_transcripts.py, and concludes with helpers/render.py, helpers/timeline_view.py, and helpers/grade.py.
  • The SKILL.md file defines the LLM reasoning logic that bridges the transcript packing and rendering stages, but execution is driven by the host agent rather than Python code.
  • No console scripts are defined in pyproject.toml, confirming that the helper scripts are the intended entry points for both developers and automated agents.

Frequently Asked Questions

Where is the main function in video-use?

There is no single main() function or __main__.py file. According to the browser-use/video-use source code, the repository is designed as a skill package loaded by an LLM agent. The practical entry points are the if __name__ == "__main__": blocks found in helpers/transcribe.py, helpers/render.py, and the other helper scripts, which allow each pipeline stage to run as a standalone command.

How do I run the video-use pipeline from the command line?

Execute the helper scripts sequentially from the repository root. Start with python helpers/transcribe.py <video_file>, then run python helpers/pack_transcripts.py <edit_dir>. After the LLM generates an EDL (using the logic in SKILL.md), run python helpers/render.py <edit_dir> to produce the final video. Each script is designed to be invoked independently without importing a central package module.

What is the purpose of the SKILL.md file?

SKILL.md serves as the specification document for the LLM agent. It contains the 12 production rules and the reasoning logic that guides the agent from the packed transcripts to the Edit-Decision-List (EDL). While it is not executable Python code, it defines the cognitive entry point for the AI-driven portion of the workflow that occurs between the packing and rendering stages.

Why are there no entry points defined in pyproject.toml?

The pyproject.toml file in the video-use repository defines package metadata and dependencies but omits console-script entry points because the tool is designed to be loaded as a skill directory by Claude Code or similar agents. The execution model relies on direct script invocation via the helpers/ modules rather than installed command-line utilities, reflecting its architecture as an agent-plugin rather than a standalone CLI tool.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →