How to Configure ElevenLabs Voice Transcription for Open Agents

To configure ElevenLabs voice transcription for Open Agents, set the ELEVENLABS_API_KEY environment variable and use the useAudioRecording hook to capture audio, which sends base-64 encoded data to the /api/transcribe route that processes the audio using the ElevenLabs Scribe v1 model.

The vercel-labs/open-agents repository ships with optional voice-input capabilities that integrate ElevenLabs speech-to-text technology. By configuring ElevenLabs voice transcription for Open Agents, you enable a "speak-to-code" experience where users can dictate commands that are transcribed before being processed by your AI agents.

Prerequisites

Before enabling voice transcription, ensure your environment meets the following requirements:

  • An active ElevenLabs API key from your ElevenLabs dashboard
  • The @ai-sdk/elevenlabs package installed in your project dependencies
  • A Next.js application environment using the App Router

Setting Up Environment Variables

The transcription service requires the ELEVENLABS_API_KEY environment variable to authenticate with the ElevenLabs API. This key is read by the ElevenLabs SDK at request time.

In apps/web/.env.example (line 30), the repository provides the configuration template:

ELEVENLABS_API_KEY=your-elevenlabs-api-key

Add this variable to your .env.local file or configure it in your Vercel Environment Variables dashboard. Without this key, the transcription endpoint will return an authentication error.

Client-Side Audio Recording

The useAudioRecording hook in apps/web/hooks/use-audio-recording.ts manages the entire audio capture workflow. According to the source code (lines 12-23 and 31-42), this hook accesses the browser microphone, converts the recorded Blob to a base-64 string, and POSTs the data to /api/transcribe.

Integrate voice recording into your React components as follows:

import { useAudioRecording } from "@/hooks/use-audio-recording";

export default function VoiceInput() {
  const {
    state,
    error,
    toggleRecording,
    clearError,
  } = useAudioRecording();

  const handleClick = async () => {
    clearError();
    const text = await toggleRecording(); // starts then stops recording
    if (text) {
      // Do something with the transcribed text, e.g. feed it to the agent
      console.log("Transcribed:", text);
    }
  };

  return (
    <button onClick={handleClick} disabled={state === "processing"}>
      {state === "recording" ? "Stop" : "Speak"}
    </button>
  );
}

The hook processes the JSON response from the server and surfaces the transcribed text to your component, handling loading states and errors automatically.

Server-Side Transcription Processing

The API route at apps/web/app/api/transcribe/route.ts receives the base-64 audio payload and invokes the transcription model. As implemented in lines 48-64, this route uses the experimental_transcribe function from the ai package with the @ai-sdk/elevenlabs provider.

import { experimental_transcribe as transcribe } from "ai";
import { elevenlabs } from "@ai-sdk/elevenlabs";

export async function POST(req: Request) {
  const { audio } = await req.json();

  const result = await transcribe({
    model: elevenlabs.transcription("scribe_v1"),
    audio,
    providerOptions: {
      elevenlabs: {
        tagAudioEvents: false,
        numSpeakers: 1,
        languageCode: "eng",
      },
    },
  });

  return Response.json({ text: result.text });
}

Provider Configuration Options

The providerOptions object allows fine-tuning of the Scribe v1 model behavior:

  • tagAudioEvents: Set to false to disable detection of non-speech sounds (such as background noise or music)
  • numSpeakers: Specifies expected speaker count (set to 1 for single-speaker input)
  • languageCode: Sets the language code (e.g., "eng" for English) to optimize transcription accuracy

The route returns { text: <transcribed-text> } on success. If the API key is missing or invalid, the ElevenLabs SDK throws an authentication error that the client handles appropriately.

Summary

  • Configure authentication: Set ELEVENLABS_API_KEY in apps/web/.env.example or your environment variables to enable the ElevenLabs SDK.
  • Capture audio on the client: Use the useAudioRecording hook from apps/web/hooks/use-audio-recording.ts to record microphone input and transmit base-64 encoded audio.
  • Process transcription on the server: The /api/transcribe route in apps/web/app/api/transcribe/route.ts handles the experimental_transcribe call using the scribe_v1 model.
  • Customize behavior: Pass providerOptions to control speaker detection, language settings, and audio event tagging for your specific use case.

Frequently Asked Questions

What audio format does the useAudioRecording hook send to the server?

The useAudioRecording hook captures audio using the browser's MediaRecorder API and converts the resulting Blob into a base-64 encoded string. This string is transmitted via POST request to the /api/transcribe endpoint, where the ElevenLabs transcription model processes the raw audio data.

How do I change the transcription language or speaker detection settings?

Modify the providerOptions.elevenlabs object in apps/web/app/api/transcribe/route.ts. Update the languageCode parameter (for example, to "spa" for Spanish) and adjust the numSpeakers value based on your audio input. These parameters are passed directly to the ElevenLabs Scribe v1 model during the experimental_transcribe invocation.

What happens if the ElevenLabs API key is missing or invalid?

If the ELEVENLABS_API_KEY environment variable is not set or contains an invalid key, the ElevenLabs SDK fails to authenticate at request time. The transcription route returns an authentication error, which the useAudioRecording hook catches and surfaces through its error state, allowing your UI to display an appropriate message to the user.

Can I use a different ElevenLabs model instead of scribe_v1?

The current implementation in apps/web/app/api/transcribe/route.ts specifically uses elevenlabs.transcription("scribe_v1") as the model parameter. To use a different model, update the model string in the experimental_transcribe configuration. Verify that your chosen model supports the audio format and provider options you are passing.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →