How to Implement DTW Gesture Learning with 3-Rehearsal Protocol on Edge Devices Using RuView
RuView provides a no-std, stack-only Dynamic Time Warping (DTW) gesture learner that implements a 3-rehearsal training protocol on ESP-32 S3 class edge devices using less than 10 KB of RAM.
RuView is an open-source Wi-Fi sensing framework that ships with a compact gesture learning module designed specifically for resource-constrained edge deployments. The DTW gesture learning with 3-rehearsal protocol enables users to teach custom gestures through three consecutive demonstrations, with the system automatically validating consistency before committing the template to flash storage.
Understanding the 3-Rehearsal Protocol in RuView
The 3-rehearsal protocol in lrn_dtw_gesture_learn.rs ensures robust gesture learning by requiring consistent motion trajectories across three separate recordings. The protocol follows this state-driven flow:
- Enter learning mode after detecting
STILLNESS_FRAMESconsecutive frames belowSTILLNESS_THRESHOLD. - Record motion trajectory when motion energy exceeds the threshold.
- Capture the trajectory when stillness returns and the buffer contains valid samples.
- Repeat steps 2-3 three times (
REHEARSALS_REQUIRED = 3). - Validate and commit if the three recordings exhibit mutual DTW distances below
LEARN_DTW_THRESHOLD, averaging them into a new template stored with IDs starting at 100.
When the learner returns to the Idle phase, the module continuously recognizes gestures by sliding a phase-delta window over incoming CSI frames and comparing against stored templates using RECOGNIZE_DTW_THRESHOLD.
Core Architecture of the DTW Gesture Learner
State Machine and Learning Phases
The LearnPhase enum in lrn_dtw_gesture_learn.rs (lines 54-65) drives the protocol through four distinct states:
- Idle: Recognition mode active, monitoring for stillness to trigger learning.
- WaitingStill: Accumulating stillness frames before allowing motion capture.
- Recording: Capturing phase-delta samples into the rehearsal buffer.
- Captured: Trajectory complete, awaiting validation against previous rehearsals.
State transitions occur based on motion energy thresholds and frame counters, ensuring deterministic behavior on edge devices without dynamic allocation.
Stillness Detection and Motion Energy
The system uses a lightweight motion-energy metric to gate learning and capture:
const STILLNESS_THRESHOLD: f32 = 0.15;
const STILLNESS_FRAMES: usize = 60; // 3 seconds at 20 Hz
The host computes motion_energy (e.g., RMS of phase deltas) and passes it to process_frame. When motion_energy < STILLNESS_THRESHOLD for STILLNESS_FRAMES consecutive calls, the learner transitions from Idle to WaitingStill, priming the system for the first rehearsal.
DTW Distance Calculation with Sakoe-Chiba Constraints
The DTW implementation uses a constrained dynamic programming approach optimized for microcontrollers:
fn dtw_distance(a: &[f32], b: &[f32]) -> f32 {
// Static 64×64 cost matrix (~4 KB stack)
// Sakoe-Chiba band width = 5
// O(N×M) where N,M ≤ 64
}
The algorithm allocates a fixed 64×64 cost matrix on the stack (approximately 4 KB), avoiding heap fragmentation. With a Sakoe-Chiba band width of 5, the computation requires roughly 4,000 floating-point operations, completing in under 10 milliseconds on the ESP-32 S3 at 240 MHz.
Template Storage and Management
The learner maintains up to MAX_TEMPLATES (16) fixed-length templates in static storage:
const MAX_TEMPLATES: usize = 16;
const TEMPLATE_LEN: usize = 64;
const REHEARSALS_REQUIRED: usize = 3;
struct Template {
data: [f32; TEMPLATE_LEN],
id: u8, // Starts at 100 to avoid built-in gesture collisions
}
Each template occupies 64 floats (256 bytes) plus metadata, totaling approximately 4.5 KB for 16 templates. User-learned gestures receive IDs starting at 100, ensuring separation from factory-defined gestures in the 0-99 range.
Implementing the Gesture Learner on Edge Devices
To integrate the DTW gesture learner into your ESP-32 S3 or similar edge deployment, instantiate the learner and feed CSI frames at 20 Hz:
use wifi_densepose_wasm_edge::lrn_dtw_gesture_learn::GestureLearner;
// Initialize the learner (stack allocation only)
let mut learner = GestureLearner::new();
// Main loop: process CSI frames at 20 Hz
fn process_csi_frame(
learner: &mut GestureLearner,
phase_deltas: &[f32], // CSI phase differences across sub-carriers
motion_energy: f32 // Computed RMS or similar metric
) {
// Process frame returns up to 4 events per call
let events = learner.process_frame(phase_deltas, motion_energy);
for (event_id, value) in events {
match event_id {
730 => handle_gesture_learned(value as u8), // New template stored
731 => handle_gesture_matched(value as u8), // Recognition event
732 => log_match_distance(value), // DTW distance of match
733 => update_template_count(value as usize), // Current template inventory
_ => {}
}
}
}
Integration requirements:
- Phase extraction: Provide the first sub-carrier's phase delta (
phases[0]) or average across carriers. - Motion energy: Compute a scalar representing movement magnitude; values below
0.15indicate stillness. - Timing: Maintain 20 Hz sampling to align with the
STILLNESS_FRAMEStiming (60 frames = 3 seconds).
Event-Driven Recognition Workflow
The learner communicates with the host application through a static event buffer emitting up to four events per frame:
| Event ID | Name | Value Meaning |
|---|---|---|
| 730 | EVENT_GESTURE_LEARNED |
Template ID of newly stored gesture (starts at 100) |
| 731 | EVENT_GESTURE_MATCHED |
Template ID of recognized gesture |
| 732 | EVENT_MATCH_DISTANCE |
Floating-point DTW distance (raw bits) of the match |
| 733 | EVENT_TEMPLATE_COUNT |
Current number of stored templates (0-16) |
These events enable reactive programming models where the host application updates UI elements, triggers actuators, or logs analytics based on gesture state changes without polling internal learner state.
Performance Characteristics on ESP-32 S3
The RuView DTW gesture learner is optimized for severe resource constraints:
- Memory footprint: < 10 KB total RAM usage (4 KB DTW matrix + 4.5 KB templates + state buffers).
- Latency: < 10 ms per DTW computation at 240 MHz (ESP-32 S3).
- Throughput: Supports continuous 20 Hz CSI processing with 16 concurrent templates.
- Storage: 16 user-defined templates (IDs 100-115) with 64-sample length each.
- WASM compatibility: Compiles to
wasm32-unknown-unknownfor sandboxed edge deployment.
The implementation avoids dynamic allocation entirely, using const generics and static arrays to ensure deterministic memory usage suitable for real-time operating systems on microcontrollers.
Summary
- RuView provides a production-ready DTW gesture learning module in
lrn_dtw_gesture_learn.rsdesigned for edge devices like the ESP-32 S3. - The 3-rehearsal protocol requires three consistent motion demonstrations validated by mutual DTW distance checks before template storage.
- The system uses constrained DTW with a static 64×64 cost matrix, completing recognition in under 10 milliseconds while consuming less than 10 KB of RAM.
- Integration requires feeding CSI phase deltas and motion energy metrics at 20 Hz, handling asynchronous events (IDs 730-733) for learning and recognition outcomes.
- All operations are no-std compatible and WASM-ready, enabling secure, sandboxed deployment on resource-constrained microcontrollers.
Frequently Asked Questions
What hardware platforms support RuView's DTW gesture learning?
RuView's DTW gesture learner targets ESP-32 S3 class microcontrollers and any platform supporting wasm32-unknown-unknown targets. The implementation requires approximately 10 KB of free RAM and a 240 MHz CPU to maintain the sub-10-millisecond DTW computation budget. The code uses no-std Rust with libm for floating-point operations, avoiding OS dependencies.
How does the 3-rehearsal protocol ensure gesture reliability?
The protocol enforces consistency by requiring three separate recordings (REHEARSALS_REQUIRED = 3) that must exhibit pairwise DTW distances below LEARN_DTW_THRESHOLD. If any rehearsal deviates significantly from the others, the learner discards the set and returns to WaitingStill for another attempt. Only mutually similar trajectories are averaged into the final 64-sample template, filtering out accidental motions.
What is the difference between learning mode and recognition mode?
Learning mode activates when the system detects sustained stillness (STILLNESS_FRAMES at 20 Hz), cycling through WaitingStill, Recording, and Captured states to collect rehearsals. During this phase, recognition is suspended. Recognition mode resumes automatically when the learner returns to Idle, continuously comparing incoming CSI phase deltas against stored templates using sliding window DTW and emitting match events (ID 731) when distances fall below RECOGNIZE_DTW_THRESHOLD.
How do I integrate the gesture learner with existing CSI data pipelines?
Integration requires instantiating GestureLearner::new() and calling process_frame(phases, motion_energy) at 20 Hz within your CSI processing loop. The phases parameter expects a slice of phase deltas (typically using the first sub-carrier or an average), while motion_energy is a scalar float representing movement magnitude (RMS of phase deltas works well). Handle the returned event tuples (IDs 730-733) to update your application state, trigger UI feedback, or log gesture analytics.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →