How VMAware's Memoization System Accelerates Repeated VM Detection

VMAware caches the boolean results and scoring data of every detection technique in a fixed-size lookup table, eliminating redundant CPUID queries, file system scans, and registry checks on subsequent calls to VM::detect().

The kernelwernel/vmaware library performs extensive low-level inspection—reading CPUID leaves, scanning firmware tables, and checking hypervisor-specific registry keys—to determine if code is running inside a virtual machine. Because these operations are deterministic and computationally expensive, VMAware implements a comprehensive memoization layer that stores technique outcomes after the first execution, ensuring repeated detection calls complete in microseconds rather than milliseconds.

The Core Memoization Architecture

The VM::memo Struct and cache_table

At the heart of the system lies the VM::memo struct defined in [src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3148). It maintains a static std::array<cache_entry, enum_size + 1> named cache_table that provides O(1) indexed storage for every detection technique. Each cache_entry records the boolean outcome, its score contribution, and an optional brand identifier, allowing the library to reconstruct the final verdict without re-executing the underlying check.

Cache Management Functions

The struct exposes four static helper methods to interact with the table: cache_store() writes results, is_cached() checks for prior execution, cache_fetch() retrieves stored data, and uncache() invalidates specific entries when necessary.

Extended Caching for Derived Data

Beyond individual technique results, VMAware memoizes auxiliary computations that would otherwise require redundant system interaction:

Integration with the Detection Pipeline

The core::run_all Execution Loop

The public API methods VM::detect() and VM::brand() delegate work to core::run_all(), implemented later in the same header. Before executing any technique, the core queries the memo table:

// Conceptual flow from core::run_all()
for (u16 flag = VM::technique_begin; flag < VM::technique_end; ++flag) {
    if (VM::memo::is_cached(flag)) {
        auto entry = VM::memo::cache_fetch(flag);
        // Accumulate score from cache, skip execution
        continue;
    }
    bool result = run_technique(flag);  // Expensive system call
    VM::memo::cache_store(flag, result, get_points(flag), get_brand(flag));
}

This logic ensures that deterministic checks are performed exactly once per process lifetime.

Performance Characteristics

The memoization layer delivers measurable speed improvements through several architectural decisions:

  • Contiguous Memory Layout – The std::array storage provides cache-friendly spatial locality and eliminates heap fragmentation.
  • System Call Elimination – CPUID instructions, file I/O, and registry queries occur only during the first detection sweep.
  • Lock-Free Access – Once populated, the cache is read-only for subsequent detections, removing synchronization overhead in multi-threaded scenarios.
  • Short-Circuit Compatibility – When the SHORTCUT flag is enabled, detection stops once a threshold score is reached; memoization ensures that partial results from an early exit remain valid for future calls.

Usage Examples

The following example demonstrates the performance differential between initial and repeated detection calls:

#include "vmaware.hpp"
#include <iostream>
#include <chrono>

int main() {
    // First call: executes all techniques, populates cache
    auto start = std::chrono::high_resolution_clock::now();
    bool is_vm = VM::detect();
    auto end = std::chrono::high_resolution_clock::now();
    std::cout << "First detection: " << is_vm 
              << " (took " << (end - start).count() << " ns)\n";

    // Second call: retrieves all results from memo::cache_table
    start = std::chrono::high_resolution_clock::now();
    is_vm = VM::detect();
    end = std::chrono::high_resolution_clock::now();
    std::cout << "Cached detection: " << is_vm 
              << " (took " << (end - start).count() << " ns)\n";

    // Brand query also leverages cached scoreboard data
    std::string brand = VM::brand();  // Instant retrieval
    std::cout << "Detected hypervisor: " << brand << "\n";
}

Key Source Files

  • src/vmaware.hpp – Contains the VM::memo struct, cache_table definition, auxiliary caches (leaf_cache, threadcount, cpu_brand), and the core::run_all() integration logic.
  • src/cli.cpp – Demonstrates real-world usage with repeated calls to VM::detect() and VM::brand(), benefiting from memoization overhead reduction.
  • README.md – High-level documentation describing the memoization behavior and performance guarantees.

Summary

  • VMAware's memoization system stores technique results in a fixed-size std::array indexed by technique enum values, enabling O(1) cache lookups.
  • The VM::memo struct caches not only boolean detection results but also brand scoreboards, CPUID leaf support bits, and hardware concurrency counts.
  • core::run_all() consults memo::is_cached() before executing any expensive check, ensuring deterministic operations run once per process.
  • Subsequent calls to VM::detect() or VM::brand() retrieve data from read-only caches, eliminating system calls and reducing execution time to microseconds.
  • The architecture is lock-free after initial population and compatible with early-exit optimization flags.

Frequently Asked Questions

What specific data does VMAware's memoization system cache?

The system caches three categories of data: (1) per-technique boolean results and scoring weights in memo::cache_table, (2) aggregated brand strings and scoreboards used by VM::brand(), and (3) auxiliary hardware information such as CPUID leaf support bits, CPU vendor strings, and thread counts. All caches reside in static storage within [src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3299).

How does memoization affect thread safety?

The memoization layer is lock-free for reads after the first detection pass. The cache_table is populated during the initial single-threaded detection sweep (or under the library's internal mutex if called concurrently), after which all entries are immutable. This design allows multiple threads to call VM::detect() simultaneously without contention, as subsequent accesses only read from the pre-computed static cache.

Can the memoization cache be cleared programmatically?

Yes. While the cache persists for the process lifetime by default, the VM::memo struct provides the uncache(flag) method to invalidate specific technique entries. This is useful when the library is used in long-running processes that need to re-evaluate the environment after major system changes, though the source code in [src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3148) primarily optimizes for the common case of immutable VM detection results.

Does memoization work with VMAware's "short-cut" detection mode?

Absolutely. When the SHORTCUT flag is enabled, detection terminates early once the confidence score exceeds the hypervisor threshold. The memoization system stores results for all techniques executed before the short-cut exit, including partial brand scoreboard data. If VM::detect() is called again, the library instantly returns the cached conclusion without re-scanning, even if the first run stopped early. This combination yields the fastest possible repeated detection path.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →