# How Graphify Handles Video and Audio Transcription with Faster-Whisper Integration

> Graphify offers fast video and audio transcription via faster-whisper integration. It downloads media, caches results, and uses knowledge graph nodes for domain-specific prompts.

- Repository: [Graphify Labs/graphify](https://github.com/Graphify-Labs/graphify)
- Tags: how-to-guide
- Published: 2026-07-15

---

**Graphify converts video and audio resources into plain-text transcripts using a faster-whisper integration that downloads remote media, caches results, and supports domain-specific prompts constructed from knowledge graph nodes.**

Graphify transforms multimedia content into structured knowledge by extracting text from video and audio sources. According to the Graphify-Labs/graphify source code, the transcription pipeline combines faster-whisper for high-performance inference with yt-dlp for media extraction, wrapped in a caching layer that avoids redundant processing.

## The Transcription Pipeline Architecture

The core implementation in [`graphify/transcribe.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/transcribe.py) orchestrates a six-step workflow that handles both local files and remote URLs.

### Step 1: URL Detection and Audio Extraction

The `is_url` function (lines 45-48) detects remote resources by checking for `http://`, `https://`, or `www.` prefixes. When a URL is detected, the `download_audio` function (lines 50-92) uses **yt-dlp** to extract the audio-only stream, caching downloaded files under `<out_dir>/downloads` for subsequent reuse.

### Step 2: Model Selection and Loading

The `_model_name` helper (lines 19-22) reads the `GRAPHIFY_WHISPER_MODEL` environment variable, defaulting to `"base"` if unset. The `_get_whisper` function (lines 23-31) initializes the faster-whisper library, raising a clear `ImportError` if the optional dependency is missing.

### Step 3: Domain-Aware Prompt Construction

Before transcription, `build_whisper_prompt` (lines 95-115) constructs an initial prompt by concatenating labels from the top "god nodes"—the most significant concepts in the knowledge graph. This cues the model about domain-specific terminology. If no nodes exist, a generic fallback prompt is used.

### Step 4: Executing Faster-Whisper

The `transcribe` function (lines 45-63) instantiates `WhisperModel` with the selected model, CPU device, and `int8` compute type for efficient inference. It processes the audio with a beam size of 5 and the constructed prompt, then concatenates segments, strips whitespace, and writes the result to a `.txt` file in the `transcripts` directory.

### Step 5: Caching and Reuse

Graphify checks `transcript_path` (lines 41-44) to determine if a transcript already exists for the given audio stem. When `force=False`, the cached file is returned immediately, eliminating re-processing overhead.

### Step 6: Batch Processing

The `transcribe_all` function (lines 66-86) iterates over lists of paths or URLs, calling `transcribe` for each entry while catching and reporting individual errors without aborting the entire batch.

## Installing the Optional Video Dependencies

The faster-whisper integration is available as an optional feature group declared in [`pyproject.toml`](https://github.com/Graphify-Labs/graphify/blob/main/pyproject.toml) (line 64). The video dependency group includes `faster-whisper` (Python ≥ 3.11) and `yt-dlp>=2026.6.9`.

Install with:

```bash
pip install "graphify[video]"

```

## Configuration Options

Environment variables control transcription behavior:

- **GRAPHIFY_WHISPER_MODEL**: Selects the model size (tiny, base, small, medium, large)
- **GRAPHIFY_WHISPER_PROMPT**: Optional manual override for the initial prompt

The implementation defaults to CPU inference with `int8` quantization, balancing speed and accuracy without requiring GPU resources.

## Practical Implementation Examples

Single local file:

```python
from pathlib import Path
from graphify.transcribe import transcribe

audio_path = Path("lecture.mp4")
transcript_path = transcribe(audio_path)
print(transcript_path.read_text())

```

Remote URL with caching:

```python
from graphify.transcribe import transcribe

url = "https://www.youtube.com/watch?v=example"
txt_path = transcribe(url)  # Downloads once, caches forever

print(txt_path)

```

Batch processing with custom model:

```python
import os
from graphify.transcribe import transcribe_all

os.environ["GRAPHIFY_WHISPER_MODEL"] = "medium"
files = ["interview.wav", "https://youtu.be/abc123"]
results = transcribe_all(files)
for path in results:
    print(f"Transcript: {path}")

```

## Summary

- Graphify's transcription pipeline in [`graphify/transcribe.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/transcribe.py) integrates faster-whisper with yt-dlp for comprehensive video and audio processing.
- The system automatically detects URLs and downloads audio streams, caching both media and transcripts to avoid redundant computation.
- Domain-specific prompts are constructed from knowledge graph "god nodes" to improve transcription accuracy for technical content.
- Configuration via `GRAPHIFY_WHISPER_MODEL` and optional dependencies installed with `pip install "graphify[video]"` provide flexibility across deployment environments.

## Frequently Asked Questions

### What model does Graphify use for transcription by default?

Graphify defaults to the `"base"` faster-whisper model unless overridden by the `GRAPHIFY_WHISPER_MODEL` environment variable. You can set this to `tiny`, `small`, `medium`, or `large` depending on your accuracy and speed requirements.

### How does Graphify handle YouTube videos and other remote URLs?

When `is_url` detects a remote resource, the `download_audio` function uses yt-dlp to extract the audio stream, saving it to `<out_dir>/downloads`. Subsequent calls with the same URL return the cached file, while the transcript itself is cached separately in the `transcripts` directory.

### What is the purpose of the "god nodes" prompt in Graphify's transcription?

The `build_whisper_prompt` function concatenates labels from the top "god nodes"—the most important concepts in your knowledge graph—to create an initial prompt for Whisper. This hints at domain-specific vocabulary, improving recognition accuracy for technical terms or specialized jargon.

### Does Graphify require a GPU for transcription?

No. According to the implementation in [`graphify/transcribe.py`](https://github.com/Graphify-Labs/graphify/blob/main/graphify/transcribe.py), the `WhisperModel` is initialized with `device="cpu"` and `compute_type="int8"`, enabling efficient transcription on CPU-only systems without requiring CUDA or other GPU accelerators.