How Dograh AudioConfig Determines Sample Rates and Buffer Sizes for the Pipeline
AudioConfig serves as the single source of truth for audio parameters in Dograh's voice AI pipeline, automatically capping pipeline sample rates at 16 kHz for VAD compatibility while deriving buffer sizes from 16-bit PCM byte requirements.
The AudioConfig class in the Dograh repository centralizes all audio-related configuration for real-time voice processing pipelines built on Pipecat. Understanding how this component determines sample rates and buffer sizes is essential for optimizing telephony integrations and ensuring consistent audio processing across transport layers.
Sample Rate Selection and Capping Logic
Transport vs. Pipeline Sample Rates
In api/services/pipecat/audio_config.py, the AudioConfig dataclass defines distinct fields for transport_in_sample_rate, transport_out_sample_rate, and pipeline_sample_rate. During instantiation, the __post_init__ method enforces a hard ceiling of 16 kHz on the pipeline sample rate to maintain compatibility with the Voice Activity Detection (VAD) component.
The logic implemented at lines 43-53 calculates:
self.pipeline_sample_rate = min(self.transport_out_sample_rate, 16000)
If the transport's output rate exceeds 16 kHz, the configuration logs a warning and caps the value, delegating resampling responsibilities to the transport layer.
VAD Rate Validation
The configuration strictly validates that vad_sample_rate is set to either 8 kHz or 16 kHz. As implemented in lines 36-41 of audio_config.py, any other value raises a ValueError immediately upon instantiation. This constraint ensures the VAD model receives audio at supported frequencies while the pipeline operates at a consistent internal rate.
Buffer Size Calculations
The configuration derives buffer dimensions from the pipeline sample rate using 16-bit PCM encoding assumptions (2 bytes per sample).
From Seconds to Bytes
The buffer_size_bytes property (lines 66-70) calculates the raw byte capacity required for the audio buffer:
buffer_size_bytes = pipeline_sample_rate * 2 * buffer_size_seconds
Here, buffer_size_seconds defaults to 5.0 seconds, representing how often the merge-audio step executes in the pipeline.
Sample Counts and Recording Limits
For components requiring sample counts rather than byte values, buffer_size_samples (lines 71-75) provides:
buffer_size_samples = pipeline_sample_rate * buffer_size_seconds
Additionally, max_recording_bytes (lines 76-80) enforces storage limits for recorded audio using the default max_recording_duration_seconds of 5 minutes:
max_recording_bytes = pipeline_sample_rate * 2 * max_recording_duration_seconds
Factory Function and Transport Integration
Provider Registry Lookup
The create_audio_config(transport_type) factory function automates sample rate configuration by consulting the telephony provider registry. Defined in lines 83-115, this function queries api/services/telephony/registry.py to retrieve a ProviderSpec containing the wire-format transport_sample_rate.
For WebRTC-based transports, the function defaults to 16 kHz when no specific provider rate is found. This ensures that even when external telephony systems operate at higher rates (e.g., 48 kHz), the internal pipeline maintains VAD-compatible parameters.
Implementation Code Examples
Direct Construction with Custom Rates
When you need explicit control over every parameter, instantiate AudioConfig directly:
from api.services.pipecat.audio_config import AudioConfig
cfg = AudioConfig(
transport_in_sample_rate=48000, # incoming from a 48 kHz telephony provider
transport_out_sample_rate=48000, # outgoing back to the provider
vad_sample_rate=16000, # VAD works at 16 kHz
pipeline_sample_rate=None, # will be auto-capped to 16 kHz
buffer_size_seconds=4.0,
)
print(cfg.pipeline_sample_rate) # → 16000 (capped)
print(cfg.buffer_size_bytes) # → 128000 bytes (16 kHz × 2 × 4 s)
print(cfg.max_recording_bytes) # → 1920000 bytes (5 min limit)
Factory-Based Configuration
For automatic provider-specific configuration, use the factory function:
from api.services.pipecat.audio_config import create_audio_config
# Retrieves the Vonage provider's wire-format sample rate from the registry
cfg = create_audio_config("vonage")
print(cfg.transport_in_sample_rate) # → e.g., 24000 (provider-specific)
print(cfg.pipeline_sample_rate) # → 16000 (auto-capped)
print(cfg.buffer_size_seconds) # → 5.0 (default)
Consuming Config in Pipeline Components
Pass the configuration to audio processing components to ensure consistent parameters:
from api.services.pipecat.audio_mixer import AudioMixer
def build_mixer(audio_cfg: AudioConfig):
mixer = AudioMixer(
sample_rate=audio_cfg.pipeline_sample_rate,
buffer_size=audio_cfg.buffer_size_samples,
)
return mixer
Summary
- 16 kHz Ceiling: The
__post_init__method capspipeline_sample_rateat 16 kHz usingmin(transport_out_sample_rate, 16000), ensuring VAD compatibility. - 16-bit PCM Assumption: Buffer calculations multiply sample rates by 2 to account for byte depth when computing
buffer_size_bytesandmax_recording_bytes. - Strict VAD Validation: The configuration rejects VAD sample rates other than 8 kHz or 16 kHz via explicit validation in the post-init hook.
- Registry Integration: The
create_audio_configfactory automatically resolves provider-specific sample rates fromapi/services/telephony/registry.py.
Frequently Asked Questions
What is the maximum pipeline sample rate supported by Dograh AudioConfig?
The maximum supported pipeline sample rate is 16 kHz. The __post_init__ method in api/services/pipecat/audio_config.py explicitly caps the value using min(self.transport_out_sample_rate, 16000), logging a warning when higher transport rates are supplied.
How is buffer_size_bytes calculated in AudioConfig?
The property calculates buffer capacity as pipeline_sample_rate * 2 * buffer_size_seconds, where the multiplier of 2 represents 16-bit PCM encoding (2 bytes per sample). This ensures the buffer can hold the complete audio data for the specified duration.
Can AudioConfig handle different sample rates for input and output transports?
Yes, the class defines separate fields for transport_in_sample_rate and transport_out_sample_rate, allowing asymmetrical configurations. However, the pipeline_sample_rate is derived exclusively from the output rate to ensure downstream VAD and mixing components receive consistently formatted audio.
How does create_audio_config determine the correct transport sample rate?
The factory function queries the telephony provider registry via registry.get_optional() to retrieve a ProviderSpec containing the provider's specific transport_sample_rate. For WebRTC transports without a registered provider, it defaults to 16 kHz.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →