How Graphify Handles Video and Audio Transcription with Faster-Whisper Integration

Graphify converts video and audio resources into plain-text transcripts using a faster-whisper integration that downloads remote media, caches results, and supports domain-specific prompts constructed from knowledge graph nodes.

Graphify transforms multimedia content into structured knowledge by extracting text from video and audio sources. According to the Graphify-Labs/graphify source code, the transcription pipeline combines faster-whisper for high-performance inference with yt-dlp for media extraction, wrapped in a caching layer that avoids redundant processing.

The Transcription Pipeline Architecture

The core implementation in graphify/transcribe.py orchestrates a six-step workflow that handles both local files and remote URLs.

Step 1: URL Detection and Audio Extraction

The is_url function (lines 45-48) detects remote resources by checking for http://, https://, or www. prefixes. When a URL is detected, the download_audio function (lines 50-92) uses yt-dlp to extract the audio-only stream, caching downloaded files under <out_dir>/downloads for subsequent reuse.

Step 2: Model Selection and Loading

The _model_name helper (lines 19-22) reads the GRAPHIFY_WHISPER_MODEL environment variable, defaulting to "base" if unset. The _get_whisper function (lines 23-31) initializes the faster-whisper library, raising a clear ImportError if the optional dependency is missing.

Step 3: Domain-Aware Prompt Construction

Before transcription, build_whisper_prompt (lines 95-115) constructs an initial prompt by concatenating labels from the top "god nodes"—the most significant concepts in the knowledge graph. This cues the model about domain-specific terminology. If no nodes exist, a generic fallback prompt is used.

Step 4: Executing Faster-Whisper

The transcribe function (lines 45-63) instantiates WhisperModel with the selected model, CPU device, and int8 compute type for efficient inference. It processes the audio with a beam size of 5 and the constructed prompt, then concatenates segments, strips whitespace, and writes the result to a .txt file in the transcripts directory.

Step 5: Caching and Reuse

Graphify checks transcript_path (lines 41-44) to determine if a transcript already exists for the given audio stem. When force=False, the cached file is returned immediately, eliminating re-processing overhead.

Step 6: Batch Processing

The transcribe_all function (lines 66-86) iterates over lists of paths or URLs, calling transcribe for each entry while catching and reporting individual errors without aborting the entire batch.

Installing the Optional Video Dependencies

The faster-whisper integration is available as an optional feature group declared in pyproject.toml (line 64). The video dependency group includes faster-whisper (Python ≥ 3.11) and yt-dlp>=2026.6.9.

Install with:

pip install "graphify[video]"

Configuration Options

Environment variables control transcription behavior:

  • GRAPHIFY_WHISPER_MODEL: Selects the model size (tiny, base, small, medium, large)
  • GRAPHIFY_WHISPER_PROMPT: Optional manual override for the initial prompt

The implementation defaults to CPU inference with int8 quantization, balancing speed and accuracy without requiring GPU resources.

Practical Implementation Examples

Single local file:

from pathlib import Path
from graphify.transcribe import transcribe

audio_path = Path("lecture.mp4")
transcript_path = transcribe(audio_path)
print(transcript_path.read_text())

Remote URL with caching:

from graphify.transcribe import transcribe

url = "https://www.youtube.com/watch?v=example"
txt_path = transcribe(url)  # Downloads once, caches forever

print(txt_path)

Batch processing with custom model:

import os
from graphify.transcribe import transcribe_all

os.environ["GRAPHIFY_WHISPER_MODEL"] = "medium"
files = ["interview.wav", "https://youtu.be/abc123"]
results = transcribe_all(files)
for path in results:
    print(f"Transcript: {path}")

Summary

  • Graphify's transcription pipeline in graphify/transcribe.py integrates faster-whisper with yt-dlp for comprehensive video and audio processing.
  • The system automatically detects URLs and downloads audio streams, caching both media and transcripts to avoid redundant computation.
  • Domain-specific prompts are constructed from knowledge graph "god nodes" to improve transcription accuracy for technical content.
  • Configuration via GRAPHIFY_WHISPER_MODEL and optional dependencies installed with pip install "graphify[video]" provide flexibility across deployment environments.

Frequently Asked Questions

What model does Graphify use for transcription by default?

Graphify defaults to the "base" faster-whisper model unless overridden by the GRAPHIFY_WHISPER_MODEL environment variable. You can set this to tiny, small, medium, or large depending on your accuracy and speed requirements.

How does Graphify handle YouTube videos and other remote URLs?

When is_url detects a remote resource, the download_audio function uses yt-dlp to extract the audio stream, saving it to <out_dir>/downloads. Subsequent calls with the same URL return the cached file, while the transcript itself is cached separately in the transcripts directory.

What is the purpose of the "god nodes" prompt in Graphify's transcription?

The build_whisper_prompt function concatenates labels from the top "god nodes"—the most important concepts in your knowledge graph—to create an initial prompt for Whisper. This hints at domain-specific vocabulary, improving recognition accuracy for technical terms or specialized jargon.

Does Graphify require a GPU for transcription?

No. According to the implementation in graphify/transcribe.py, the WhisperModel is initialized with device="cpu" and compute_type="int8", enabling efficient transcription on CPU-only systems without requiring CUDA or other GPU accelerators.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →