How to Find the Main Entry Point and Core Modules in the video-use Repository
The video-use repository does not expose a traditional console script entry point; instead, execution flows through six standalone helper scripts in the helpers/ directory, each containing an if __name__ == "__main__": block that serves as the practical entry point for the video editing pipeline.
The video-use skill is a Python package designed to run inside a Claude-Code (or other LLM-driven) agent for automated video editing. Unlike conventional Python packages that define entry points in setup.py or pyproject.toml, this repository operates as a collection of stage-specific utilities that the host agent invokes sequentially. Understanding where the execution begins requires examining the helper scripts that handle transcription, packing, rendering, and grading.
Understanding the Repository Architecture
The repository follows a pipeline-based architecture rather than a monolithic application structure. According to the browser-use/video-use source code, the package is loaded by the host agent via a directory link at ~/.claude/skills/video-use, which means there is no single main() function to import. Instead, the workflow is divided into discrete stages, each implemented by a dedicated helper script that can be executed independently from the command line.
The core execution flow progresses through these stages:
- Transcribe – Extracts audio and generates JSON transcripts using the ElevenLabs Scribe API
- Pack Transcripts – Consolidates transcripts into a single
takes_packed.mdfile for LLM consumption - Reasoning/Planning – The LLM processes the packed transcripts and generates an Edit-Decision-List (EDL) as specified in
SKILL.md - Render – Applies the EDL to produce the final video using
ffmpeg - Self-Evaluation – Generates preview frames for visual verification
- Grading – Evaluates output quality and determines if re-rendering is necessary (max 3 attempts)
The Pipeline-Based Entry Points
Each stage in the editing pipeline corresponds to a specific file in the helpers/ directory. These files represent the core modules of the repository and contain the executable logic for the video processing workflow.
Transcribe Stage
The helpers/transcribe.py file serves as the typical starting point for the pipeline. It extracts audio from raw video files, sends the data to the ElevenLabs Scribe API, and caches the resulting JSON transcripts in the edit/transcripts/ directory.
# From helpers/transcribe.py
if __name__ == "__main__":
# Handles single video transcription
# Creates: edit/transcripts/{video_name}.json
For batch processing, helpers/transcribe_batch.py extends this functionality to handle multiple videos in a directory.
Pack Transcripts Stage
Once transcription is complete, helpers/pack_transcripts.py consolidates all individual JSON transcripts into a single markdown file named takes_packed.md. This file (approximately 12KB) is the primary input that the LLM reads to understand the video content and make editing decisions.
python helpers/pack_transcripts.py ./edit
Render Stage
The helpers/render.py module constitutes the primary output stage, turning the LLM-generated EDL into the final video file. It executes ffmpeg commands to apply cuts, transitions, and optional animation overlays, producing edit/final.mp4 along with intermediate files in edit/render/.
python helpers/render.py ./edit
Evaluation and Grading
The final stages utilize helpers/timeline_view.py for generating diagnostic PNG images of specific time ranges (used during self-evaluation loops) and helpers/grade.py for scoring the rendered output. The grading module determines whether the video meets quality thresholds or requires re-rendering (up to 3 attempts).
# Generate timeline visualization for verification
python helpers/timeline_view.py ./edit 12.3 15.7
How to Locate Entry Points in the Source Code
To programmatically identify all executable modules in the repository, search for the standard Python entry point pattern:
rg "if __name__ == \"__main__\"" -g "*.py"
This search returns six helper files that serve as the practical main entry points:
helpers/transcribe.pyhelpers/transcribe_batch.pyhelpers/pack_transcripts.pyhelpers/render.pyhelpers/timeline_view.pyhelpers/grade.py
Additionally, examining pyproject.toml confirms that no console-script entry points are defined, verifying that execution is manual via the helper scripts. The README.md and SKILL.md files provide the high-level workflow documentation that points developers to these scripts as the canonical entry points.
Practical Usage Examples
Transcribing a Single Video
Start the pipeline by processing an individual raw take:
python helpers/transcribe.py path/to/raw_take.mp4
Result: Creates edit/transcripts/raw_take.json and caches the extracted audio file.
Packing Transcripts for LLM Analysis
Consolidate all transcripts before invoking the LLM reasoning stage:
python helpers/pack_transcripts.py ./edit
Result: Generates edit/takes_packed.md ready for the LLM agent.
Rendering the Final Video
After the LLM produces an EDL (stored in the edit directory), generate the final output:
python helpers/render.py ./edit
Result: Produces edit/final.mp4 with all edits applied.
Running a Complete Batch Workflow
Process multiple videos through the entire pipeline:
# Transcribe all videos in directory
python helpers/transcribe_batch.py ./videos
# Prepare consolidated transcript
python helpers/pack_transcripts.py ./videos/edit
# LLM reasoning happens in the host agent (Claude Code) using SKILL.md
# Render final output
python helpers/render.py ./videos/edit
# Evaluate quality
python helpers/grade.py ./videos/edit
Summary
- The video-use repository lacks a traditional single entry point; instead, it distributes functionality across six helper scripts in the
helpers/directory. - Each core module contains an
if __name__ == "__main__":block, making them directly executable from the command line. - The standard workflow begins with
helpers/transcribe.py, proceeds throughhelpers/pack_transcripts.py, and concludes withhelpers/render.py,helpers/timeline_view.py, andhelpers/grade.py. - The
SKILL.mdfile defines the LLM reasoning logic that bridges the transcript packing and rendering stages, but execution is driven by the host agent rather than Python code. - No console scripts are defined in
pyproject.toml, confirming that the helper scripts are the intended entry points for both developers and automated agents.
Frequently Asked Questions
Where is the main function in video-use?
There is no single main() function or __main__.py file. According to the browser-use/video-use source code, the repository is designed as a skill package loaded by an LLM agent. The practical entry points are the if __name__ == "__main__": blocks found in helpers/transcribe.py, helpers/render.py, and the other helper scripts, which allow each pipeline stage to run as a standalone command.
How do I run the video-use pipeline from the command line?
Execute the helper scripts sequentially from the repository root. Start with python helpers/transcribe.py <video_file>, then run python helpers/pack_transcripts.py <edit_dir>. After the LLM generates an EDL (using the logic in SKILL.md), run python helpers/render.py <edit_dir> to produce the final video. Each script is designed to be invoked independently without importing a central package module.
What is the purpose of the SKILL.md file?
SKILL.md serves as the specification document for the LLM agent. It contains the 12 production rules and the reasoning logic that guides the agent from the packed transcripts to the Edit-Decision-List (EDL). While it is not executable Python code, it defines the cognitive entry point for the AI-driven portion of the workflow that occurs between the packing and rendering stages.
Why are there no entry points defined in pyproject.toml?
The pyproject.toml file in the video-use repository defines package metadata and dependencies but omits console-script entry points because the tool is designed to be loaded as a skill directory by Claude Code or similar agents. The execution model relies on direct script invocation via the helpers/ modules rather than installed command-line utilities, reflecting its architecture as an agent-plugin rather than a standalone CLI tool.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →