How camofox-browser Extracts YouTube Transcripts Using yt-dlp with Browser Fallback

camofox-browser automatically routes YouTube transcript requests through yt-dlp for speed, then falls back to Playwright browser automation when the binary is missing or captions are dynamically generated.

The camofox-browser repository provides a resilient YouTube transcript extraction pipeline via a single HTTP endpoint. This open-source tool combines binary execution with headless Firefox automation to handle both static subtitle files and live caption streams.

The Two-Stage Extraction Architecture

The implementation follows a deterministic two-stage strategy defined in server.js. It prioritizes the yt-dlp binary for performance, then automatically escalates to browser-based extraction for edge cases.

Stage 1: yt-dlp Binary Detection and Execution

At startup, the module scans for yt-dlp in common locations such as yt-dlp and /usr/local/bin/yt-dlp. The detection logic in lib/youtube.js (lines 88-99) caches the valid path in ytDlpPath for reuse.

When a POST /youtube/transcript request arrives, the handler invokes ensureYtDlp() to lazily re-detect the binary if initial detection failed (see server.js lines 74-76). If present, ytDlpTranscript() (lines 15-99) executes:

  • Normalizes the video URL and requested language
  • Creates a temporary directory for downloads
  • Runs yt-dlp twice: first for video metadata, then for subtitle extraction in json3, vtt, or srv3 format
  • Selects the first successfully written subtitle file
  • Parses content using parseJson3, parseVtt, or parseXml (defined in lib/youtube.js lines 5-30)
  • Returns a standardized JSON payload containing status, transcript, video_url, video_title, language, and word count

This approach avoids the overhead of browser startup when subtitles are available for direct download.

Stage 2: Playwright Browser Fallback

If yt-dlp is absent or returns errors (e.g., "no captions"), the handler falls back to browserTranscript() in server.js (lines 106-166). This function uses the same headless Firefox engine that powers the rest of camofox-browser.

The browser fallback executes these steps:

  • Opens a fresh browser context named __yt_transcript__ dedicated to extraction
  • Mutes the video element to prevent audio playback
  • Listens for network responses containing the api/timedtext endpoint
  • Attempts Strategy A: Extracts the caption track URL baked into ytInitialPlayerResponse (most reliable)
  • Falls back to Strategy B: Plays the video and captures timed-text responses as they arrive
  • Parses the raw payload using the same three parsers (parseJson3, parseVtt, parseXml) as the yt-dlp path

This ensures uniform output format regardless of extraction method.

Proxy Configuration

Both extraction paths honor rotating proxy settings configured via the PROXY_POOL_URL environment variable. The buildProxyUrl function in lib/proxy.js (lines 244-254) constructs the proxy URL at startup.

The computed URL passes to ytDlpTranscript() as the proxyUrl argument (injected as --proxy <url>) and to Playwright page instances when browser fallback activates. This prevents IP blocking during high-volume extraction.

Practical Usage Examples

Basic curl Request

Request transcripts using the automated fallback pipeline:

curl -X POST https://your-camofox-host:9377/youtube/transcript \
     -H "Content-Type: application/json" \
     -d '{"url":"https://www.youtube.com/watch?v=dQw4w9WgXcQ","languages":["en"]}'

Successful responses return structured JSON:

{
  "status": "ok",
  "transcript": "[00:00] We're no strangers to love\n[00:04] You know the rules ...",
  "video_url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "video_id": "dQw4w9WgXcQ",
  "video_title": "Never Gonna Give You Up",
  "language": "en",
  "total_words": 423,
  "available_languages": [
    {"code":"en","name":"English","kind":"auto"},
    {"code":"es","name":"Español","kind":"manual"}
  ]
}

Node.js Integration

Integrate transcript extraction into your applications:

import fetch from 'node-fetch';

async function getTranscript(videoUrl) {
  const res = await fetch('http://localhost:9377/youtube/transcript', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ url: videoUrl, languages: ['en'] })
  });

  const data = await res.json();
  if (data.status !== 'ok') throw new Error(data.message);
  console.log('Transcript:', data.transcript);
}

await getTranscript('https://youtu.be/dQw4w9WgXcQ');

Enabling Proxy Support

Set the environment variable before starting the server:

export PROXY_POOL_URL="http://my-proxy:3128"
npm start

No code changes are required; the server automatically applies the proxy to both yt-dlp and browser contexts.

Summary

  • camofox-browser provides a single POST /youtube/transcript endpoint that abstracts yt-dlp and browser extraction complexity
  • Stage 1 uses ytDlpTranscript() in lib/youtube.js to download subtitles directly via binary execution when available
  • Stage 2 falls back to browserTranscript() in server.js using Playwright to capture live captions from api/timedtext endpoints
  • Both paths support proxy configuration via PROXY_POOL_URL and output identical JSON structures containing transcript text, metadata, and available languages
  • The unified parser trio (parseJson3, parseVtt, parseXml) ensures consistent formatting across binary and browser extraction methods

Frequently Asked Questions

What happens if yt-dlp is not installed?

If the binary is missing, the ensureYtDlp() function returns null, triggering the browserTranscript() fallback automatically. The system attempts Strategy A (parsing ytInitialPlayerResponse) first, then Strategy B (capturing network responses during playback) to retrieve captions without requiring yt-dlp installation.

Which subtitle formats does camofox-browser support?

The implementation supports JSON3, VTT, and SRV3/XML formats. The ytDlpTranscript() function requests these formats explicitly, and browserTranscript() parses the live caption stream using the same three parsers defined in lib/youtube.js to maintain output consistency.

How does the proxy configuration apply to both extraction methods?

The buildProxyUrl function in lib/proxy.js constructs the proxy URL at server startup. This URL passes to ytDlpTranscript() as a command-line --proxy argument and to Playwright browser contexts, ensuring both yt-dlp binary calls and browser automation route through the specified proxy pool.

Why does the browser fallback use a dedicated context named __yt_transcript__?

The __yt_transcript__ context isolates transcript extraction from other browser operations, preventing cookie contamination and resource conflicts. This dedicated context configuration also allows specific optimizations like video muting and network interception rules that apply only to caption extraction workflows.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →