# How to Implement DTW Gesture Learning with 3-Rehearsal Protocol on Edge Devices Using RuView

> Implement DTW gesture learning with 3-rehearsal protocol on edge devices using RuView. RuView offers a no-std, stack-only DTW learner under 10KB RAM for ESP-32 S3.

- Repository: [rUv/RuView](https://github.com/ruvnet/RuView)
- Tags: how-to-guide
- Published: 2026-03-08

---

**RuView provides a no-std, stack-only Dynamic Time Warping (DTW) gesture learner that implements a 3-rehearsal training protocol on ESP-32 S3 class edge devices using less than 10 KB of RAM.**

RuView is an open-source Wi-Fi sensing framework that ships with a compact gesture learning module designed specifically for resource-constrained edge deployments. The **DTW gesture learning with 3-rehearsal protocol** enables users to teach custom gestures through three consecutive demonstrations, with the system automatically validating consistency before committing the template to flash storage.

## Understanding the 3-Rehearsal Protocol in RuView

The 3-rehearsal protocol in [`lrn_dtw_gesture_learn.rs`](https://github.com/ruvnet/RuView/blob/main/lrn_dtw_gesture_learn.rs) ensures robust gesture learning by requiring consistent motion trajectories across three separate recordings. The protocol follows this state-driven flow:

1. **Enter learning mode** after detecting `STILLNESS_FRAMES` consecutive frames below `STILLNESS_THRESHOLD`.
2. **Record motion trajectory** when motion energy exceeds the threshold.
3. **Capture the trajectory** when stillness returns and the buffer contains valid samples.
4. **Repeat steps 2-3 three times** (`REHEARSALS_REQUIRED = 3`).
5. **Validate and commit** if the three recordings exhibit mutual DTW distances below `LEARN_DTW_THRESHOLD`, averaging them into a new template stored with IDs starting at 100.

When the learner returns to the `Idle` phase, the module continuously recognizes gestures by sliding a phase-delta window over incoming CSI frames and comparing against stored templates using `RECOGNIZE_DTW_THRESHOLD`.

## Core Architecture of the DTW Gesture Learner

### State Machine and Learning Phases

The `LearnPhase` enum in [`lrn_dtw_gesture_learn.rs`](https://github.com/ruvnet/RuView/blob/main/lrn_dtw_gesture_learn.rs) (lines 54-65) drives the protocol through four distinct states:

- **Idle**: Recognition mode active, monitoring for stillness to trigger learning.
- **WaitingStill**: Accumulating stillness frames before allowing motion capture.
- **Recording**: Capturing phase-delta samples into the rehearsal buffer.
- **Captured**: Trajectory complete, awaiting validation against previous rehearsals.

State transitions occur based on motion energy thresholds and frame counters, ensuring deterministic behavior on edge devices without dynamic allocation.

### Stillness Detection and Motion Energy

The system uses a lightweight motion-energy metric to gate learning and capture:

```rust
const STILLNESS_THRESHOLD: f32 = 0.15;
const STILLNESS_FRAMES: usize = 60; // 3 seconds at 20 Hz

```

The host computes `motion_energy` (e.g., RMS of phase deltas) and passes it to `process_frame`. When `motion_energy < STILLNESS_THRESHOLD` for `STILLNESS_FRAMES` consecutive calls, the learner transitions from `Idle` to `WaitingStill`, priming the system for the first rehearsal.

### DTW Distance Calculation with Sakoe-Chiba Constraints

The DTW implementation uses a constrained dynamic programming approach optimized for microcontrollers:

```rust
fn dtw_distance(a: &[f32], b: &[f32]) -> f32 {
    // Static 64×64 cost matrix (~4 KB stack)
    // Sakoe-Chiba band width = 5
    // O(N×M) where N,M ≤ 64
}

```

The algorithm allocates a fixed 64×64 cost matrix on the stack (approximately 4 KB), avoiding heap fragmentation. With a Sakoe-Chiba band width of 5, the computation requires roughly 4,000 floating-point operations, completing in under 10 milliseconds on the ESP-32 S3 at 240 MHz.

### Template Storage and Management

The learner maintains up to `MAX_TEMPLATES` (16) fixed-length templates in static storage:

```rust
const MAX_TEMPLATES: usize = 16;
const TEMPLATE_LEN: usize = 64;
const REHEARSALS_REQUIRED: usize = 3;

struct Template {
    data: [f32; TEMPLATE_LEN],
    id: u8, // Starts at 100 to avoid built-in gesture collisions
}

```

Each template occupies 64 floats (256 bytes) plus metadata, totaling approximately 4.5 KB for 16 templates. User-learned gestures receive IDs starting at 100, ensuring separation from factory-defined gestures in the 0-99 range.

## Implementing the Gesture Learner on Edge Devices

To integrate the DTW gesture learner into your ESP-32 S3 or similar edge deployment, instantiate the learner and feed CSI frames at 20 Hz:

```rust
use wifi_densepose_wasm_edge::lrn_dtw_gesture_learn::GestureLearner;

// Initialize the learner (stack allocation only)
let mut learner = GestureLearner::new();

// Main loop: process CSI frames at 20 Hz
fn process_csi_frame(
    learner: &mut GestureLearner,
    phase_deltas: &[f32],  // CSI phase differences across sub-carriers
    motion_energy: f32     // Computed RMS or similar metric
) {
    // Process frame returns up to 4 events per call
    let events = learner.process_frame(phase_deltas, motion_energy);
    
    for (event_id, value) in events {
        match event_id {
            730 => handle_gesture_learned(value as u8),    // New template stored
            731 => handle_gesture_matched(value as u8),     // Recognition event
            732 => log_match_distance(value),               // DTW distance of match
            733 => update_template_count(value as usize),  // Current template inventory
            _ => {}
        }
    }
}

```

**Integration requirements:**

- **Phase extraction**: Provide the first sub-carrier's phase delta (`phases[0]`) or average across carriers.
- **Motion energy**: Compute a scalar representing movement magnitude; values below `0.15` indicate stillness.
- **Timing**: Maintain 20 Hz sampling to align with the `STILLNESS_FRAMES` timing (60 frames = 3 seconds).

## Event-Driven Recognition Workflow

The learner communicates with the host application through a static event buffer emitting up to four events per frame:

| Event ID | Name | Value Meaning |
|----------|------|---------------|
| **730** | `EVENT_GESTURE_LEARNED` | Template ID of newly stored gesture (starts at 100) |
| **731** | `EVENT_GESTURE_MATCHED` | Template ID of recognized gesture |
| **732** | `EVENT_MATCH_DISTANCE` | Floating-point DTW distance (raw bits) of the match |
| **733** | `EVENT_TEMPLATE_COUNT` | Current number of stored templates (0-16) |

These events enable reactive programming models where the host application updates UI elements, triggers actuators, or logs analytics based on gesture state changes without polling internal learner state.

## Performance Characteristics on ESP-32 S3

The RuView DTW gesture learner is optimized for severe resource constraints:

- **Memory footprint**: < 10 KB total RAM usage (4 KB DTW matrix + 4.5 KB templates + state buffers).
- **Latency**: < 10 ms per DTW computation at 240 MHz (ESP-32 S3).
- **Throughput**: Supports continuous 20 Hz CSI processing with 16 concurrent templates.
- **Storage**: 16 user-defined templates (IDs 100-115) with 64-sample length each.
- **WASM compatibility**: Compiles to `wasm32-unknown-unknown` for sandboxed edge deployment.

The implementation avoids dynamic allocation entirely, using const generics and static arrays to ensure deterministic memory usage suitable for real-time operating systems on microcontrollers.

## Summary

- **RuView** provides a production-ready **DTW gesture learning** module in [`lrn_dtw_gesture_learn.rs`](https://github.com/ruvnet/RuView/blob/main/lrn_dtw_gesture_learn.rs) designed for **edge devices** like the ESP-32 S3.
- The **3-rehearsal protocol** requires three consistent motion demonstrations validated by mutual DTW distance checks before template storage.
- The system uses **constrained DTW** with a static 64×64 cost matrix, completing recognition in under 10 milliseconds while consuming less than 10 KB of RAM.
- Integration requires feeding **CSI phase deltas** and **motion energy metrics** at 20 Hz, handling asynchronous events (IDs 730-733) for learning and recognition outcomes.
- All operations are **no-std compatible** and **WASM-ready**, enabling secure, sandboxed deployment on resource-constrained microcontrollers.

## Frequently Asked Questions

### What hardware platforms support RuView's DTW gesture learning?

RuView's DTW gesture learner targets **ESP-32 S3** class microcontrollers and any platform supporting `wasm32-unknown-unknown` targets. The implementation requires approximately 10 KB of free RAM and a 240 MHz CPU to maintain the sub-10-millisecond DTW computation budget. The code uses no-std Rust with `libm` for floating-point operations, avoiding OS dependencies.

### How does the 3-rehearsal protocol ensure gesture reliability?

The protocol enforces consistency by requiring **three separate recordings** (`REHEARSALS_REQUIRED = 3`) that must exhibit pairwise DTW distances below `LEARN_DTW_THRESHOLD`. If any rehearsal deviates significantly from the others, the learner discards the set and returns to `WaitingStill` for another attempt. Only mutually similar trajectories are averaged into the final 64-sample template, filtering out accidental motions.

### What is the difference between learning mode and recognition mode?

**Learning mode** activates when the system detects sustained stillness (`STILLNESS_FRAMES` at 20 Hz), cycling through `WaitingStill`, `Recording`, and `Captured` states to collect rehearsals. During this phase, recognition is suspended. **Recognition mode** resumes automatically when the learner returns to `Idle`, continuously comparing incoming CSI phase deltas against stored templates using sliding window DTW and emitting match events (ID 731) when distances fall below `RECOGNIZE_DTW_THRESHOLD`.

### How do I integrate the gesture learner with existing CSI data pipelines?

Integration requires instantiating `GestureLearner::new()` and calling `process_frame(phases, motion_energy)` at 20 Hz within your CSI processing loop. The `phases` parameter expects a slice of phase deltas (typically using the first sub-carrier or an average), while `motion_energy` is a scalar float representing movement magnitude (RMS of phase deltas works well). Handle the returned event tuples (IDs 730-733) to update your application state, trigger UI feedback, or log gesture analytics.