# How camofox-browser Extracts YouTube Transcripts Using yt-dlp with Browser Fallback

> Learn how camofox-browser extracts YouTube transcripts using yt-dlp with a Playwright browser fallback for missing binaries or dynamic captions. Get fast, reliable transcriptions.

- Repository: [jo/camofox-browser](https://github.com/jo-inc/camofox-browser)
- Tags: how-to-guide
- Published: 2026-04-15

---

**camofox-browser automatically routes YouTube transcript requests through yt-dlp for speed, then falls back to Playwright browser automation when the binary is missing or captions are dynamically generated.**

The **camofox-browser** repository provides a resilient **YouTube transcript extraction** pipeline via a single HTTP endpoint. This open-source tool combines binary execution with headless Firefox automation to handle both static subtitle files and live caption streams.

## The Two-Stage Extraction Architecture

The implementation follows a deterministic two-stage strategy defined in [`server.js`](https://github.com/jo-inc/camofox-browser/blob/main/server.js). It prioritizes the yt-dlp binary for performance, then automatically escalates to browser-based extraction for edge cases.

### Stage 1: yt-dlp Binary Detection and Execution

At startup, the module scans for **yt-dlp** in common locations such as `yt-dlp` and `/usr/local/bin/yt-dlp`. The detection logic in **[`lib/youtube.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/youtube.js)** (lines 88-99) caches the valid path in `ytDlpPath` for reuse.

When a `POST /youtube/transcript` request arrives, the handler invokes **`ensureYtDlp()`** to lazily re-detect the binary if initial detection failed (see [`server.js`](https://github.com/jo-inc/camofox-browser/blob/main/server.js) lines 74-76). If present, **`ytDlpTranscript()`** (lines 15-99) executes:

- Normalizes the video URL and requested language
- Creates a temporary directory for downloads
- Runs yt-dlp twice: first for video metadata, then for subtitle extraction in `json3`, `vtt`, or `srv3` format
- Selects the first successfully written subtitle file
- Parses content using `parseJson3`, `parseVtt`, or `parseXml` (defined in [`lib/youtube.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/youtube.js) lines 5-30)
- Returns a standardized JSON payload containing `status`, `transcript`, `video_url`, `video_title`, `language`, and word count

This approach avoids the overhead of browser startup when subtitles are available for direct download.

### Stage 2: Playwright Browser Fallback

If yt-dlp is absent or returns errors (e.g., "no captions"), the handler falls back to **`browserTranscript()`** in [`server.js`](https://github.com/jo-inc/camofox-browser/blob/main/server.js) (lines 106-166). This function uses the same headless Firefox engine that powers the rest of camofox-browser.

The browser fallback executes these steps:

- Opens a fresh browser context named `__yt_transcript__` dedicated to extraction
- Mutes the video element to prevent audio playback
- Listens for network responses containing the `api/timedtext` endpoint
- Attempts Strategy A: Extracts the caption track URL baked into `ytInitialPlayerResponse` (most reliable)
- Falls back to Strategy B: Plays the video and captures timed-text responses as they arrive
- Parses the raw payload using the same three parsers (`parseJson3`, `parseVtt`, `parseXml`) as the yt-dlp path

This ensures uniform output format regardless of extraction method.

## Proxy Configuration

Both extraction paths honor rotating proxy settings configured via the `PROXY_POOL_URL` environment variable. The **`buildProxyUrl`** function in **[`lib/proxy.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/proxy.js)** (lines 244-254) constructs the proxy URL at startup.

The computed URL passes to `ytDlpTranscript()` as the `proxyUrl` argument (injected as `--proxy <url>`) and to Playwright page instances when browser fallback activates. This prevents IP blocking during high-volume extraction.

## Practical Usage Examples

### Basic curl Request

Request transcripts using the automated fallback pipeline:

```bash
curl -X POST https://your-camofox-host:9377/youtube/transcript \
     -H "Content-Type: application/json" \
     -d '{"url":"https://www.youtube.com/watch?v=dQw4w9WgXcQ","languages":["en"]}'

```

Successful responses return structured JSON:

```json
{
  "status": "ok",
  "transcript": "[00:00] We're no strangers to love\n[00:04] You know the rules ...",
  "video_url": "https://www.youtube.com/watch?v=dQw4w9WgXcQ",
  "video_id": "dQw4w9WgXcQ",
  "video_title": "Never Gonna Give You Up",
  "language": "en",
  "total_words": 423,
  "available_languages": [
    {"code":"en","name":"English","kind":"auto"},
    {"code":"es","name":"Español","kind":"manual"}
  ]
}

```

### Node.js Integration

Integrate transcript extraction into your applications:

```javascript
import fetch from 'node-fetch';

async function getTranscript(videoUrl) {
  const res = await fetch('http://localhost:9377/youtube/transcript', {
    method: 'POST',
    headers: { 'Content-Type': 'application/json' },
    body: JSON.stringify({ url: videoUrl, languages: ['en'] })
  });

  const data = await res.json();
  if (data.status !== 'ok') throw new Error(data.message);
  console.log('Transcript:', data.transcript);
}

await getTranscript('https://youtu.be/dQw4w9WgXcQ');

```

### Enabling Proxy Support

Set the environment variable before starting the server:

```bash
export PROXY_POOL_URL="http://my-proxy:3128"
npm start

```

No code changes are required; the server automatically applies the proxy to both yt-dlp and browser contexts.

## Summary

- **camofox-browser** provides a single `POST /youtube/transcript` endpoint that abstracts yt-dlp and browser extraction complexity
- **Stage 1** uses `ytDlpTranscript()` in [`lib/youtube.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/youtube.js) to download subtitles directly via binary execution when available
- **Stage 2** falls back to `browserTranscript()` in [`server.js`](https://github.com/jo-inc/camofox-browser/blob/main/server.js) using Playwright to capture live captions from `api/timedtext` endpoints
- Both paths support **proxy configuration** via `PROXY_POOL_URL` and output identical JSON structures containing transcript text, metadata, and available languages
- The unified parser trio (`parseJson3`, `parseVtt`, `parseXml`) ensures consistent formatting across binary and browser extraction methods

## Frequently Asked Questions

### What happens if yt-dlp is not installed?

If the binary is missing, the `ensureYtDlp()` function returns null, triggering the `browserTranscript()` fallback automatically. The system attempts Strategy A (parsing `ytInitialPlayerResponse`) first, then Strategy B (capturing network responses during playback) to retrieve captions without requiring yt-dlp installation.

### Which subtitle formats does camofox-browser support?

The implementation supports **JSON3**, **VTT**, and **SRV3/XML** formats. The `ytDlpTranscript()` function requests these formats explicitly, and `browserTranscript()` parses the live caption stream using the same three parsers defined in [`lib/youtube.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/youtube.js) to maintain output consistency.

### How does the proxy configuration apply to both extraction methods?

The `buildProxyUrl` function in [`lib/proxy.js`](https://github.com/jo-inc/camofox-browser/blob/main/lib/proxy.js) constructs the proxy URL at server startup. This URL passes to `ytDlpTranscript()` as a command-line `--proxy` argument and to Playwright browser contexts, ensuring both yt-dlp binary calls and browser automation route through the specified proxy pool.

### Why does the browser fallback use a dedicated context named `__yt_transcript__`?

The `__yt_transcript__` context isolates transcript extraction from other browser operations, preventing cookie contamination and resource conflicts. This dedicated context configuration also allows specific optimizations like video muting and network interception rules that apply only to caption extraction workflows.