How to Configure Azure Services for Voice-Pro: Complete Setup Guide

Voice-Pro automatically switches from free Edge-TTS and Deep-Translator fallbacks to Azure Cognitive Services when you populate the required environment variables in a .env file.

Voice-Pro supports Azure Speech Service and Azure Translator Text API for enterprise-grade speech synthesis and translation. This guide explains how to configure Azure services for Voice-Pro by setting environment variables that the application detects at runtime. All configuration logic resides in the abus-aikorea/voice-pro repository's Python modules.

Where Azure Configuration Lives in Voice-Pro

Voice-Pro uses a modular architecture to detect and initialize Azure services. The following components handle Azure configuration and dynamic backend selection:

  • app/abus_config.py – Contains helper functions get_azure_speech_key(), get_azure_speech_region(), get_azure_translator_key(), get_azure_translator_endpoint(), and get_azure_translator_region() that read values from the environment.

  • app/abus_genuine.py – Implements azure_text_api_working(), which returns True only when all required Azure environment variables are present.

  • app/abus_tts_azure.py – Defines the AzureTTS class that sends synthesis requests to Azure Speech Service.

  • app/abus_translate_azure.py – Defines the AzureTranslator class that calls the Azure Translator Text API.

  • UI Controllers (app/gradio_tts_edge.py, app/gradio_gulliver.py) – Dynamically select Azure implementations when available: self.tts = AzureTTS() if azure_text_api_working() else EdgeTTS().

When azure_text_api_working() returns True, the UI displays Azure-TTS instead of Edge-TTS and routes all downstream pipelines through Azure-backed services.

Required Environment Variables

Create a .env file in the repository root (copy from .env.example). Populate the following entries with values obtained from your Azure portal:

Variable Description Example Value
AZURE_SPEECH_KEY Azure Speech Service subscription key abcd1234efgh5678ijkl9012mnop3456
AZURE_SPEECH_REGION Region identifier (e.g., eastus) eastus
AZURE_TRANSLATOR_KEY Azure Translator Text subscription key qrst1234uvwx5678yzab9012cdef3456
AZURE_TRANSLATOR_ENDPOINT Translator endpoint URL https://api.cognitive.microsofttranslator.com
AZURE_TRANSLATOR_REGION Region for Translator (often same as Speech) eastus

The .env.example file contains placeholders and comments explaining each variable.

Step-by-Step Azure Configuration Guide

Follow these steps to configure Azure services for Voice-Pro:

  1. Obtain Azure Keys – In the Azure portal, create a Speech resource and a Translator resource. Copy the subscription keys and region names from the "Keys and Endpoint" section of each resource.

  2. Create .env File – Duplicate the repository's .env.example to .env. Replace placeholder strings with the keys and regions from step 1.

  3. Verify Configuration Load – Run python start-voice.py voice (or use start.bat/start.sh). The console logs "Using Azure Translator API" when the variables are detected correctly in abus_genuine.py.

  4. Select Azure in UI – In the TTS tab, confirm the label displays Azure-TTS instead of Edge-TTS. All generated speech now synthesizes via Azure's service.

How Voice-Pro Detects and Uses Azure Services

The application uses runtime detection to switch between free and Azure backends. The azure_text_api_working() function in app/abus_genuine.py validates the configuration:


# app/abus_genuine.py – helper that decides whether Azure is enabled

def azure_text_api_working() -> bool:
    # Returns True only when all required env vars are present

    return (
        get_azure_speech_key() is not None
        and get_azure_speech_region() is not None
        and get_azure_translator_key() is not None
        and get_azure_translator_endpoint() is not None
    )

UI controllers dynamically instantiate the appropriate TTS class based on this check:


# app/gradio_tts_edge.py – dynamic TTS backend selection

class GradioTTSEdge:
    def __init__(self):
        # If Azure is configured, use it; otherwise fall back to Edge‑TTS

        self.tts = AzureTTS() if azure_text_api_working() else EdgeTTS()
        self.translator = (
            AzureTranslator() if azure_text_api_working() else DeepTranslator()
        )

The AzureTTS class in app/abus_tts_azure.py constructs the API endpoint and sends SSML payloads:


# app/abus_tts_azure.py – minimal Azure TTS request

class AzureTTS:
    def __init__(self):
        self.key = get_azure_speech_key()
        self.region = get_azure_speech_region()
        self.endpoint = f"https://{self.region}.tts.speech.microsoft.com/cognitiveservices/v1"

    def synthesize(self, text: str, voice: str = "en-US-JennyNeural") -> bytes:
        headers = {
            "Ocp-Apim-Subscription-Key": self.key,
            "Content-Type": "application/ssml+xml",
        }
        ssml = f"""<speak version='1.0' xml:lang='en-US'>
                     <voice name='{voice}'>{text}</voice></speak>"""
        resp = requests.post(self.endpoint, data=ssml.encode("utf-8"), headers=headers)
        resp.raise_for_status()
        return resp.content

Summary

  • Voice-Pro uses environment variables to detect and configure Azure Cognitive Services automatically.
  • Set AZURE_SPEECH_KEY, AZURE_SPEECH_REGION, AZURE_TRANSLATOR_KEY, AZURE_TRANSLATOR_ENDPOINT, and AZURE_TRANSLATOR_REGION in a .env file.
  • The azure_text_api_working() function in app/abus_genuine.py validates configuration and triggers the Azure backend.
  • UI controllers in gradio_*.py files switch between Azure-TTS and Edge-TTS based on environment variable presence.
  • Azure services provide higher-quality cloud-based speech synthesis and translation compared to free fallbacks.

Frequently Asked Questions

What Azure resources do I need to create for Voice-Pro?

You need to create two separate resources in the Azure portal: an Azure Speech Service resource for text-to-speech functionality and an Azure Translator resource for translation services. Each resource provides its own subscription key and region identifier that you must add to your .env file.

How does Voice-Pro switch between Azure and free services?

Voice-Pro calls the azure_text_api_working() function at startup to check for the presence of all required Azure environment variables. If the variables exist, the application instantiates AzureTTS and AzureTranslator classes; otherwise, it falls back to EdgeTTS and DeepTranslator classes. This logic is implemented in app/abus_genuine.py and the various gradio_*.py controller files.

Can I use different regions for Azure Speech and Azure Translator?

Yes, you can specify different regions for each service using the AZURE_SPEECH_REGION and AZURE_TRANSLATOR_REGION variables in your .env file. However, for optimal latency and cost efficiency, it is recommended to deploy both resources in the same region (such as eastus or westeurope).

Where should I place the .env file for Voice-Pro to detect it?

Place the .env file in the repository root directory, at the same level as start-voice.py. Voice-Pro loads environment variables from this file at runtime, using the template provided in .env.example as a reference for the required variable names and formats.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →