How to Configure ElevenLabs Voice Transcription for Open Agents
To configure ElevenLabs voice transcription for Open Agents, set the ELEVENLABS_API_KEY environment variable and use the useAudioRecording hook to capture audio, which sends base-64 encoded data to the /api/transcribe route that processes the audio using the ElevenLabs Scribe v1 model.
The vercel-labs/open-agents repository ships with optional voice-input capabilities that integrate ElevenLabs speech-to-text technology. By configuring ElevenLabs voice transcription for Open Agents, you enable a "speak-to-code" experience where users can dictate commands that are transcribed before being processed by your AI agents.
Prerequisites
Before enabling voice transcription, ensure your environment meets the following requirements:
- An active ElevenLabs API key from your ElevenLabs dashboard
- The
@ai-sdk/elevenlabspackage installed in your project dependencies - A Next.js application environment using the App Router
Setting Up Environment Variables
The transcription service requires the ELEVENLABS_API_KEY environment variable to authenticate with the ElevenLabs API. This key is read by the ElevenLabs SDK at request time.
In apps/web/.env.example (line 30), the repository provides the configuration template:
ELEVENLABS_API_KEY=your-elevenlabs-api-key
Add this variable to your .env.local file or configure it in your Vercel Environment Variables dashboard. Without this key, the transcription endpoint will return an authentication error.
Client-Side Audio Recording
The useAudioRecording hook in apps/web/hooks/use-audio-recording.ts manages the entire audio capture workflow. According to the source code (lines 12-23 and 31-42), this hook accesses the browser microphone, converts the recorded Blob to a base-64 string, and POSTs the data to /api/transcribe.
Integrate voice recording into your React components as follows:
import { useAudioRecording } from "@/hooks/use-audio-recording";
export default function VoiceInput() {
const {
state,
error,
toggleRecording,
clearError,
} = useAudioRecording();
const handleClick = async () => {
clearError();
const text = await toggleRecording(); // starts then stops recording
if (text) {
// Do something with the transcribed text, e.g. feed it to the agent
console.log("Transcribed:", text);
}
};
return (
<button onClick={handleClick} disabled={state === "processing"}>
{state === "recording" ? "Stop" : "Speak"}
</button>
);
}
The hook processes the JSON response from the server and surfaces the transcribed text to your component, handling loading states and errors automatically.
Server-Side Transcription Processing
The API route at apps/web/app/api/transcribe/route.ts receives the base-64 audio payload and invokes the transcription model. As implemented in lines 48-64, this route uses the experimental_transcribe function from the ai package with the @ai-sdk/elevenlabs provider.
import { experimental_transcribe as transcribe } from "ai";
import { elevenlabs } from "@ai-sdk/elevenlabs";
export async function POST(req: Request) {
const { audio } = await req.json();
const result = await transcribe({
model: elevenlabs.transcription("scribe_v1"),
audio,
providerOptions: {
elevenlabs: {
tagAudioEvents: false,
numSpeakers: 1,
languageCode: "eng",
},
},
});
return Response.json({ text: result.text });
}
Provider Configuration Options
The providerOptions object allows fine-tuning of the Scribe v1 model behavior:
- tagAudioEvents: Set to
falseto disable detection of non-speech sounds (such as background noise or music) - numSpeakers: Specifies expected speaker count (set to
1for single-speaker input) - languageCode: Sets the language code (e.g.,
"eng"for English) to optimize transcription accuracy
The route returns { text: <transcribed-text> } on success. If the API key is missing or invalid, the ElevenLabs SDK throws an authentication error that the client handles appropriately.
Summary
- Configure authentication: Set
ELEVENLABS_API_KEYinapps/web/.env.exampleor your environment variables to enable the ElevenLabs SDK. - Capture audio on the client: Use the
useAudioRecordinghook fromapps/web/hooks/use-audio-recording.tsto record microphone input and transmit base-64 encoded audio. - Process transcription on the server: The
/api/transcriberoute inapps/web/app/api/transcribe/route.tshandles theexperimental_transcribecall using thescribe_v1model. - Customize behavior: Pass
providerOptionsto control speaker detection, language settings, and audio event tagging for your specific use case.
Frequently Asked Questions
What audio format does the useAudioRecording hook send to the server?
The useAudioRecording hook captures audio using the browser's MediaRecorder API and converts the resulting Blob into a base-64 encoded string. This string is transmitted via POST request to the /api/transcribe endpoint, where the ElevenLabs transcription model processes the raw audio data.
How do I change the transcription language or speaker detection settings?
Modify the providerOptions.elevenlabs object in apps/web/app/api/transcribe/route.ts. Update the languageCode parameter (for example, to "spa" for Spanish) and adjust the numSpeakers value based on your audio input. These parameters are passed directly to the ElevenLabs Scribe v1 model during the experimental_transcribe invocation.
What happens if the ElevenLabs API key is missing or invalid?
If the ELEVENLABS_API_KEY environment variable is not set or contains an invalid key, the ElevenLabs SDK fails to authenticate at request time. The transcription route returns an authentication error, which the useAudioRecording hook catches and surfaces through its error state, allowing your UI to display an appropriate message to the user.
Can I use a different ElevenLabs model instead of scribe_v1?
The current implementation in apps/web/app/api/transcribe/route.ts specifically uses elevenlabs.transcription("scribe_v1") as the model parameter. To use a different model, update the model string in the experimental_transcribe configuration. Verify that your chosen model supports the audio format and provider options you are passing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →