# How to Configure ElevenLabs Voice Transcription for Open Agents

> Learn to configure ElevenLabs voice transcription for Open Agents. Set your API key and use the useAudioRecording hook to capture and transcribe audio with ElevenLabs Scribe v1.

- Repository: [Vercel Labs/open-agents](https://github.com/vercel-labs/open-agents)
- Tags: how-to-guide
- Published: 2026-04-16

---

**To configure ElevenLabs voice transcription for Open Agents, set the `ELEVENLABS_API_KEY` environment variable and use the `useAudioRecording` hook to capture audio, which sends base-64 encoded data to the `/api/transcribe` route that processes the audio using the ElevenLabs Scribe v1 model.**

The vercel-labs/open-agents repository ships with optional voice-input capabilities that integrate ElevenLabs speech-to-text technology. By configuring ElevenLabs voice transcription for Open Agents, you enable a "speak-to-code" experience where users can dictate commands that are transcribed before being processed by your AI agents.

## Prerequisites

Before enabling voice transcription, ensure your environment meets the following requirements:

- An active **ElevenLabs API key** from your ElevenLabs dashboard
- The `@ai-sdk/elevenlabs` package installed in your project dependencies
- A Next.js application environment using the App Router

## Setting Up Environment Variables

The transcription service requires the `ELEVENLABS_API_KEY` environment variable to authenticate with the ElevenLabs API. This key is read by the ElevenLabs SDK at request time.

In `apps/web/.env.example` (line 30), the repository provides the configuration template:

```dotenv
ELEVENLABS_API_KEY=your-elevenlabs-api-key

```

Add this variable to your `.env.local` file or configure it in your Vercel Environment Variables dashboard. Without this key, the transcription endpoint will return an authentication error.

## Client-Side Audio Recording

The **`useAudioRecording`** hook in [`apps/web/hooks/use-audio-recording.ts`](https://github.com/vercel-labs/open-agents/blob/main/apps/web/hooks/use-audio-recording.ts) manages the entire audio capture workflow. According to the source code (lines 12-23 and 31-42), this hook accesses the browser microphone, converts the recorded Blob to a base-64 string, and POSTs the data to `/api/transcribe`.

Integrate voice recording into your React components as follows:

```tsx
import { useAudioRecording } from "@/hooks/use-audio-recording";

export default function VoiceInput() {
  const {
    state,
    error,
    toggleRecording,
    clearError,
  } = useAudioRecording();

  const handleClick = async () => {
    clearError();
    const text = await toggleRecording(); // starts then stops recording
    if (text) {
      // Do something with the transcribed text, e.g. feed it to the agent
      console.log("Transcribed:", text);
    }
  };

  return (
    <button onClick={handleClick} disabled={state === "processing"}>
      {state === "recording" ? "Stop" : "Speak"}
    </button>
  );
}

```

The hook processes the JSON response from the server and surfaces the transcribed text to your component, handling loading states and errors automatically.

## Server-Side Transcription Processing

The API route at [`apps/web/app/api/transcribe/route.ts`](https://github.com/vercel-labs/open-agents/blob/main/apps/web/app/api/transcribe/route.ts) receives the base-64 audio payload and invokes the transcription model. As implemented in lines 48-64, this route uses the **`experimental_transcribe`** function from the `ai` package with the `@ai-sdk/elevenlabs` provider.

```typescript
import { experimental_transcribe as transcribe } from "ai";
import { elevenlabs } from "@ai-sdk/elevenlabs";

export async function POST(req: Request) {
  const { audio } = await req.json();

  const result = await transcribe({
    model: elevenlabs.transcription("scribe_v1"),
    audio,
    providerOptions: {
      elevenlabs: {
        tagAudioEvents: false,
        numSpeakers: 1,
        languageCode: "eng",
      },
    },
  });

  return Response.json({ text: result.text });
}

```

### Provider Configuration Options

The **`providerOptions`** object allows fine-tuning of the **Scribe v1** model behavior:

- **tagAudioEvents**: Set to `false` to disable detection of non-speech sounds (such as background noise or music)
- **numSpeakers**: Specifies expected speaker count (set to `1` for single-speaker input)
- **languageCode**: Sets the language code (e.g., `"eng"` for English) to optimize transcription accuracy

The route returns `{ text: <transcribed-text> }` on success. If the API key is missing or invalid, the ElevenLabs SDK throws an authentication error that the client handles appropriately.

## Summary

- **Configure authentication**: Set `ELEVENLABS_API_KEY` in `apps/web/.env.example` or your environment variables to enable the ElevenLabs SDK.
- **Capture audio on the client**: Use the `useAudioRecording` hook from [`apps/web/hooks/use-audio-recording.ts`](https://github.com/vercel-labs/open-agents/blob/main/apps/web/hooks/use-audio-recording.ts) to record microphone input and transmit base-64 encoded audio.
- **Process transcription on the server**: The `/api/transcribe` route in [`apps/web/app/api/transcribe/route.ts`](https://github.com/vercel-labs/open-agents/blob/main/apps/web/app/api/transcribe/route.ts) handles the `experimental_transcribe` call using the `scribe_v1` model.
- **Customize behavior**: Pass `providerOptions` to control speaker detection, language settings, and audio event tagging for your specific use case.

## Frequently Asked Questions

### What audio format does the useAudioRecording hook send to the server?

The `useAudioRecording` hook captures audio using the browser's MediaRecorder API and converts the resulting Blob into a base-64 encoded string. This string is transmitted via POST request to the `/api/transcribe` endpoint, where the ElevenLabs transcription model processes the raw audio data.

### How do I change the transcription language or speaker detection settings?

Modify the `providerOptions.elevenlabs` object in [`apps/web/app/api/transcribe/route.ts`](https://github.com/vercel-labs/open-agents/blob/main/apps/web/app/api/transcribe/route.ts). Update the `languageCode` parameter (for example, to `"spa"` for Spanish) and adjust the `numSpeakers` value based on your audio input. These parameters are passed directly to the ElevenLabs Scribe v1 model during the `experimental_transcribe` invocation.

### What happens if the ElevenLabs API key is missing or invalid?

If the `ELEVENLABS_API_KEY` environment variable is not set or contains an invalid key, the ElevenLabs SDK fails to authenticate at request time. The transcription route returns an authentication error, which the `useAudioRecording` hook catches and surfaces through its `error` state, allowing your UI to display an appropriate message to the user.

### Can I use a different ElevenLabs model instead of scribe_v1?

The current implementation in [`apps/web/app/api/transcribe/route.ts`](https://github.com/vercel-labs/open-agents/blob/main/apps/web/app/api/transcribe/route.ts) specifically uses `elevenlabs.transcription("scribe_v1")` as the model parameter. To use a different model, update the model string in the `experimental_transcribe` configuration. Verify that your chosen model supports the audio format and provider options you are passing.