# How System Audio Capture Works on macOS (ScreenCaptureKit) vs Windows (WASAPI) in Meetily

> Discover how Meetily captures system audio on macOS with ScreenCaptureKit and Windows with WASAPI. Understand the core technologies driving cross-platform audio.

- Repository: [Zackriya Solutions/meetily](https://github.com/Zackriya-Solutions/meetily)
- Tags: internals
- Published: 2026-08-01

---

**Meetily implements cross-platform system audio capture using Apple's ScreenCaptureKit on macOS and Microsoft's WASAPI on Windows, both wrapped in CPAL streams for unified pipeline processing.**

Meetily, an open-source AI meeting assistant by Zackriya-Solutions, routes operating system audio directly into its transcription pipeline through platform-specific native APIs. This article breaks down how **macOS ScreenCaptureKit** and **Windows WASAPI** capture implementations differ under the hood, with reference to the actual source code in the meetily repository.

## ScreenCaptureKit on macOS: Direct Digital Audio Capture

Meetily leverages **ScreenCaptureKit** as the default system audio capture backend on macOS. This framework provides hardware-accelerated access to the digital audio stream before it reaches any physical output device.

### How ScreenCaptureKit Delivers Audio

In [`frontend/src-tauri/src/audio/capture/backend_config.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/capture/backend_config.rs), the `AudioCaptureBackend` enum defines `ScreenCaptureKit` with the description "Apple's ScreenCaptureKit framework – Higher level API with good compatibility". The backend grants access to the pristine pulse-code modulation (PCM) stream—the exact data your Mac sends to speakers or headphones—unaffected by Bluetooth codec compression or analog conversion losses.

The implementation in [`frontend/src-tauri/src/audio/system_audio_stream.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/system_audio_stream.rs) creates a CPAL stream using the ScreenCaptureKit host. When initialization succeeds, the code logs the event and begins feeding audio frames into the processing pipeline. A comment in [`frontend/src-tauri/src/audio/devices/speakers.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/devices/speakers.rs) confirms this approach enforces "consistent sample rates" across capture sessions.

### macOS Fallback Behavior

If ScreenCaptureKit initialization fails—common on macOS versions prior to Monterey 12.3—the pipeline automatically falls back to the **CoreAudio** host. This resilience ensures recordings proceed even when the preferred API is unavailable.

## WASAPI on Windows: Loopback Capture Architecture

On Windows, Meetily selects **WASAPI** (Windows Audio Session API) as its system audio capture mechanism. WASAPI operates in **loopback mode**, capturing the final mixed output stream that Windows sends to the selected playback device.

### WASAPI Host Initialization

The Windows-specific device enumeration lives in [`frontend/src-tauri/src/audio/devices/platform/windows.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/devices/platform/windows.rs). Here, Meetily instantiates a WASAPI host through CPAL before querying available endpoints. The code handles initialization failures gracefully, logging diagnostic messages when the host cannot be created.

### Device Type Detection via Naming Patterns

WASAPI exposes device characteristics through human-readable names. Meetily parses these in [`frontend/src-tauri/src/audio/device_detection.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/device_detection.rs) to categorize hardware:

```rust
// audio/device_detection.rs
if device_name.contains("Bluetooth") {
    info!("✅ Windows WASAPI: Bluetooth Audio prefix detected for '{}'", device_name);
}

```

This pattern matching distinguishes **Bluetooth**, **USB**, and **built-in** audio devices without requiring additional Windows API calls.

### WASAPI Stream Lifecycle

Before recording termination in [`recording_manager.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/recording_manager.rs), Meetily explicitly stops the device monitor to halt WASAPI polling. This ordered shutdown prevents audio glitches and resource leaks during rapid start/stop cycles.

## Platform-Specific Backend Selection

Meetily abstracts platform differences through a unified backend configuration layer. The default backend selection uses conditional compilation:

```rust
// audio/capture/backend_config.rs
pub enum AudioCaptureBackend {
    ScreenCaptureKit,
    WASAPI,
    CoreAudio,
    // …
}

impl AudioCaptureBackend {
    pub fn default() -> Self {
        #[cfg(target_os = "macos")]
        return AudioCaptureBackend::ScreenCaptureKit;
        #[cfg(target_os = "windows")]
        return AudioCaptureBackend::Wasapi;
    }
}

```

### Stream Construction Pipeline

Both backends feed into the same CPAL abstraction layer in [`audio/stream.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/audio/stream.rs):

```rust
let host = match backend_type {
    AudioCaptureBackend::ScreenCaptureKit => {
        cpal::host_from_id(cpal::HostId::ScreenCaptureKit)?
    }
    AudioCaptureBackend::WASAPI => {
        cpal::host_from_id(cpal::HostId::Wasapi)?
    }
    _ => cpal::default_host(),
};

let device = host.default_output_device().ok_or("No output device")?;
let config = device.default_output_config()?;

```

The resulting stream delivers **48 kHz PCM audio** regardless of platform, enabling fixed-size window processing later in the pipeline.

## Integration with Meetily's Audio Pipeline

Captured system audio flows through three processing stages before reaching Whisper transcription:

1. **Stream creation** – `audio::capture::system::create_stream()` instantiates the platform-appropriate backend
2. **Mixing and VAD** – [`audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/audio/pipeline.rs) combines system audio with microphone input, then applies voice activity detection filtering
3. **Transcription forwarding** – processed frames arrive at the Whisper engine as normalized 16 kHz or 24 kHz audio (resampled from the 48 kHz capture rate)

This consistent sample rate contract between ScreenCaptureKit and WASAPI backends simplifies downstream audio mathematics.

### Frontend Invocation

TypeScript code in the Meetily UI triggers platform-appropriate capture:

```typescript
await invoke('start_recording', {
  mic_device_name: 'Built-in Microphone',
  system_device_name: 'ScreenCaptureKit', // macOS
  // system_device_name: 'WASAPI',        // Windows
  meeting_name: 'Team Sync'
});

```

## Key Implementation Files

| File | Responsibility |
|------|--------------|
| [`frontend/src-tauri/src/audio/capture/backend_config.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/capture/backend_config.rs) | Backend enum definitions and default selection logic |
| [`frontend/src-tauri/src/audio/devices/platform/macos.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/devices/platform/macos.rs) | macOS ScreenCaptureKit device enumeration |
| [`frontend/src-tauri/src/audio/devices/platform/windows.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/devices/platform/windows.rs) | Windows WASAPI device enumeration |
| [`frontend/src-tauri/src/audio/system_audio_stream.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/system_audio_stream.rs) | Stream creation and fallback handling |
| [`frontend/src-tauri/src/audio/pipeline.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/pipeline.rs) | Audio mixing, VAD, and transcription routing |
| [`frontend/src-tauri/src/audio/device_detection.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/frontend/src-tauri/src/audio/device_detection.rs) | WASAPI device type classification |

## Summary

- **macOS ScreenCaptureKit** captures digital audio before output processing, delivering high-fidelity streams with consistent sample rates
- **Windows WASAPI** operates in loopback mode, capturing the mixed output stream across diverse hardware types
- Both backends normalize to **48 kHz PCM** through CPAL, enabling unified downstream processing
- Platform selection occurs at compile time via `#[cfg(target_os = "...")]` attributes in [`backend_config.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/backend_config.rs)
- Fallback mechanisms ensure recording continuity when primary backends fail

## Frequently Asked Questions

### Does ScreenCaptureKit capture audio from specific applications or the entire system?

ScreenCaptureKit captures the **entire system audio mix**—all applications producing sound simultaneously. Meetily does not currently implement per-application audio isolation, though ScreenCaptureKit's `SCStreamConfiguration` API theoretically supports this filtering for future enhancement.

### Why doesn't Meetily use PulseAudio or PipeWire on Linux?

Meetily's Linux implementation follows a similar backend abstraction pattern but targets **ALSA** as the default capture mechanism. The [`backend_config.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/backend_config.rs) file reserves expansion space for PipeWire hosts, mirroring the macOS/Windows conditional compilation approach.

### Can WASAPI capture separate audio channels or only stereo mix?

The current Meetily implementation captures the **standard stereo mix** that Windows presents to playback devices. WASAPI supports raw channel access through `WAVEFORMATEXTENSIBLE` formats, but this would require modifications to the CPAL host integration in [`system_audio_stream.rs`](https://github.com/Zackriya-Solutions/meetily/blob/main/system_audio_stream.rs).

### What happens if both ScreenCaptureKit and CoreAudio fail on macOS?

Meetily logs the failure chain and returns an error to the frontend, which surfaces a user-facing notification. The recording manager prevents pipeline initialization when no viable capture backend exists, avoiding silent audio failures.