# How VMAware's Memoization System Accelerates Repeated VM Detection

> Discover how VMAware's memoization system speeds up repeated VM detection by caching results and eliminating redundant checks. Learn more about optimizing performance.

- Repository: [Louis/vmaware](https://github.com/kernelwernel/vmaware)
- Tags: performance
- Published: 2026-03-05

---

**VMAware caches the boolean results and scoring data of every detection technique in a fixed-size lookup table, eliminating redundant CPUID queries, file system scans, and registry checks on subsequent calls to `VM::detect()`.**

The [kernelwernel/vmaware](https://github.com/kernelwernel/vmaware) library performs extensive low-level inspection—reading CPUID leaves, scanning firmware tables, and checking hypervisor-specific registry keys—to determine if code is running inside a virtual machine. Because these operations are **deterministic** and computationally expensive, VMAware implements a comprehensive **memoization** layer that stores technique outcomes after the first execution, ensuring repeated detection calls complete in microseconds rather than milliseconds.

## The Core Memoization Architecture

### The VM::memo Struct and cache_table

At the heart of the system lies the `VM::memo` struct defined in [[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3148). It maintains a static `std::array<cache_entry, enum_size + 1>` named `cache_table` that provides O(1) indexed storage for every detection technique. Each `cache_entry` records the boolean outcome, its score contribution, and an optional brand identifier, allowing the library to reconstruct the final verdict without re-executing the underlying check.

### Cache Management Functions

The struct exposes four static helper methods to interact with the table: `cache_store()` writes results, `is_cached()` checks for prior execution, `cache_fetch()` retrieves stored data, and `uncache()` invalidates specific entries when necessary.

## Extended Caching for Derived Data

Beyond individual technique results, VMAware memoizes auxiliary computations that would otherwise require redundant system interaction:

- **Brand Resolution** – The `single_brand` and `multi_brand` strings are cached after the first call to `VM::brand()` ([[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3175-L3190)), avoiding repeated scoreboard traversal.
- **Brand Scoreboard** – The internal `brand_list` vector, which tracks hypervisor scores during detection, is preserved across calls ([[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3192-L3205)).
- **CPUID Leaf Support** – The `leaf_cache` stores support bits for specific CPUID leaves, preventing redundant processor capability checks ([[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3277-L3299)).
- **CPU Vendor String** – The `cpu_brand` cache holds the CPU vendor identifier retrieved via CPUID ([[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3736-L3745)).
- **Thread Count** – The `threadcount` cache lazily stores `std::thread::hardware_concurrency()` after its first invocation ([[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3248-L3256)).

## Integration with the Detection Pipeline

### The core::run_all Execution Loop

The public API methods `VM::detect()` and `VM::brand()` delegate work to `core::run_all()`, implemented later in the same header. Before executing any technique, the core queries the memo table:

```cpp
// Conceptual flow from core::run_all()
for (u16 flag = VM::technique_begin; flag < VM::technique_end; ++flag) {
    if (VM::memo::is_cached(flag)) {
        auto entry = VM::memo::cache_fetch(flag);
        // Accumulate score from cache, skip execution
        continue;
    }
    bool result = run_technique(flag);  // Expensive system call
    VM::memo::cache_store(flag, result, get_points(flag), get_brand(flag));
}

```

This logic ensures that deterministic checks are performed exactly once per process lifetime.

## Performance Characteristics

The memoization layer delivers measurable speed improvements through several architectural decisions:

- **Contiguous Memory Layout** – The `std::array` storage provides cache-friendly spatial locality and eliminates heap fragmentation.
- **System Call Elimination** – CPUID instructions, file I/O, and registry queries occur only during the first detection sweep.
- **Lock-Free Access** – Once populated, the cache is read-only for subsequent detections, removing synchronization overhead in multi-threaded scenarios.
- **Short-Circuit Compatibility** – When the `SHORTCUT` flag is enabled, detection stops once a threshold score is reached; memoization ensures that partial results from an early exit remain valid for future calls.

## Usage Examples

The following example demonstrates the performance differential between initial and repeated detection calls:

```cpp
#include "vmaware.hpp"
#include <iostream>
#include <chrono>

int main() {
    // First call: executes all techniques, populates cache
    auto start = std::chrono::high_resolution_clock::now();
    bool is_vm = VM::detect();
    auto end = std::chrono::high_resolution_clock::now();
    std::cout << "First detection: " << is_vm 
              << " (took " << (end - start).count() << " ns)\n";

    // Second call: retrieves all results from memo::cache_table
    start = std::chrono::high_resolution_clock::now();
    is_vm = VM::detect();
    end = std::chrono::high_resolution_clock::now();
    std::cout << "Cached detection: " << is_vm 
              << " (took " << (end - start).count() << " ns)\n";

    // Brand query also leverages cached scoreboard data
    std::string brand = VM::brand();  // Instant retrieval
    std::cout << "Detected hypervisor: " << brand << "\n";
}

```

## Key Source Files

- **[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)** – Contains the `VM::memo` struct, `cache_table` definition, auxiliary caches (`leaf_cache`, `threadcount`, `cpu_brand`), and the `core::run_all()` integration logic.
- **[`src/cli.cpp`](https://github.com/kernelwernel/vmaware/blob/main/src/cli.cpp)** – Demonstrates real-world usage with repeated calls to `VM::detect()` and `VM::brand()`, benefiting from memoization overhead reduction.
- **[`README.md`](https://github.com/kernelwernel/vmaware/blob/main/README.md)** – High-level documentation describing the memoization behavior and performance guarantees.

## Summary

- VMAware's **memoization system** stores technique results in a fixed-size `std::array` indexed by technique enum values, enabling O(1) cache lookups.
- The `VM::memo` struct caches not only boolean detection results but also brand scoreboards, CPUID leaf support bits, and hardware concurrency counts.
- **`core::run_all()`** consults `memo::is_cached()` before executing any expensive check, ensuring deterministic operations run once per process.
- Subsequent calls to **`VM::detect()`** or **`VM::brand()`** retrieve data from read-only caches, eliminating system calls and reducing execution time to microseconds.
- The architecture is **lock-free** after initial population and compatible with early-exit optimization flags.

## Frequently Asked Questions

### What specific data does VMAware's memoization system cache?

The system caches three categories of data: (1) per-technique boolean results and scoring weights in `memo::cache_table`, (2) aggregated brand strings and scoreboards used by `VM::brand()`, and (3) auxiliary hardware information such as CPUID leaf support bits, CPU vendor strings, and thread counts. All caches reside in static storage within [[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3299).

### How does memoization affect thread safety?

The memoization layer is **lock-free** for reads after the first detection pass. The `cache_table` is populated during the initial single-threaded detection sweep (or under the library's internal mutex if called concurrently), after which all entries are immutable. This design allows multiple threads to call `VM::detect()` simultaneously without contention, as subsequent accesses only read from the pre-computed static cache.

### Can the memoization cache be cleared programmatically?

Yes. While the cache persists for the process lifetime by default, the `VM::memo` struct provides the `uncache(flag)` method to invalidate specific technique entries. This is useful when the library is used in long-running processes that need to re-evaluate the environment after major system changes, though the source code in [[`src/vmaware.hpp`](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp)](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3148) primarily optimizes for the common case of immutable VM detection results.

### Does memoization work with VMAware's "short-cut" detection mode?

Absolutely. When the `SHORTCUT` flag is enabled, detection terminates early once the confidence score exceeds the hypervisor threshold. The memoization system stores results for all techniques executed before the short-cut exit, including partial brand scoreboard data. If `VM::detect()` is called again, the library instantly returns the cached conclusion without re-scanning, even if the first run stopped early. This combination yields the fastest possible repeated detection path.