How VMAware's Memoization System Accelerates Repeated VM Detection
VMAware caches the boolean results and scoring data of every detection technique in a fixed-size lookup table, eliminating redundant CPUID queries, file system scans, and registry checks on subsequent calls to VM::detect().
The kernelwernel/vmaware library performs extensive low-level inspection—reading CPUID leaves, scanning firmware tables, and checking hypervisor-specific registry keys—to determine if code is running inside a virtual machine. Because these operations are deterministic and computationally expensive, VMAware implements a comprehensive memoization layer that stores technique outcomes after the first execution, ensuring repeated detection calls complete in microseconds rather than milliseconds.
The Core Memoization Architecture
The VM::memo Struct and cache_table
At the heart of the system lies the VM::memo struct defined in [src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3148). It maintains a static std::array<cache_entry, enum_size + 1> named cache_table that provides O(1) indexed storage for every detection technique. Each cache_entry records the boolean outcome, its score contribution, and an optional brand identifier, allowing the library to reconstruct the final verdict without re-executing the underlying check.
Cache Management Functions
The struct exposes four static helper methods to interact with the table: cache_store() writes results, is_cached() checks for prior execution, cache_fetch() retrieves stored data, and uncache() invalidates specific entries when necessary.
Extended Caching for Derived Data
Beyond individual technique results, VMAware memoizes auxiliary computations that would otherwise require redundant system interaction:
- Brand Resolution – The
single_brandandmulti_brandstrings are cached after the first call toVM::brand()([src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3175-L3190)), avoiding repeated scoreboard traversal. - Brand Scoreboard – The internal
brand_listvector, which tracks hypervisor scores during detection, is preserved across calls ([src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3192-L3205)). - CPUID Leaf Support – The
leaf_cachestores support bits for specific CPUID leaves, preventing redundant processor capability checks ([src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3277-L3299)). - CPU Vendor String – The
cpu_brandcache holds the CPU vendor identifier retrieved via CPUID ([src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3736-L3745)). - Thread Count – The
threadcountcache lazily storesstd::thread::hardware_concurrency()after its first invocation ([src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3248-L3256)).
Integration with the Detection Pipeline
The core::run_all Execution Loop
The public API methods VM::detect() and VM::brand() delegate work to core::run_all(), implemented later in the same header. Before executing any technique, the core queries the memo table:
// Conceptual flow from core::run_all()
for (u16 flag = VM::technique_begin; flag < VM::technique_end; ++flag) {
if (VM::memo::is_cached(flag)) {
auto entry = VM::memo::cache_fetch(flag);
// Accumulate score from cache, skip execution
continue;
}
bool result = run_technique(flag); // Expensive system call
VM::memo::cache_store(flag, result, get_points(flag), get_brand(flag));
}
This logic ensures that deterministic checks are performed exactly once per process lifetime.
Performance Characteristics
The memoization layer delivers measurable speed improvements through several architectural decisions:
- Contiguous Memory Layout – The
std::arraystorage provides cache-friendly spatial locality and eliminates heap fragmentation. - System Call Elimination – CPUID instructions, file I/O, and registry queries occur only during the first detection sweep.
- Lock-Free Access – Once populated, the cache is read-only for subsequent detections, removing synchronization overhead in multi-threaded scenarios.
- Short-Circuit Compatibility – When the
SHORTCUTflag is enabled, detection stops once a threshold score is reached; memoization ensures that partial results from an early exit remain valid for future calls.
Usage Examples
The following example demonstrates the performance differential between initial and repeated detection calls:
#include "vmaware.hpp"
#include <iostream>
#include <chrono>
int main() {
// First call: executes all techniques, populates cache
auto start = std::chrono::high_resolution_clock::now();
bool is_vm = VM::detect();
auto end = std::chrono::high_resolution_clock::now();
std::cout << "First detection: " << is_vm
<< " (took " << (end - start).count() << " ns)\n";
// Second call: retrieves all results from memo::cache_table
start = std::chrono::high_resolution_clock::now();
is_vm = VM::detect();
end = std::chrono::high_resolution_clock::now();
std::cout << "Cached detection: " << is_vm
<< " (took " << (end - start).count() << " ns)\n";
// Brand query also leverages cached scoreboard data
std::string brand = VM::brand(); // Instant retrieval
std::cout << "Detected hypervisor: " << brand << "\n";
}
Key Source Files
src/vmaware.hpp– Contains theVM::memostruct,cache_tabledefinition, auxiliary caches (leaf_cache,threadcount,cpu_brand), and thecore::run_all()integration logic.src/cli.cpp– Demonstrates real-world usage with repeated calls toVM::detect()andVM::brand(), benefiting from memoization overhead reduction.README.md– High-level documentation describing the memoization behavior and performance guarantees.
Summary
- VMAware's memoization system stores technique results in a fixed-size
std::arrayindexed by technique enum values, enabling O(1) cache lookups. - The
VM::memostruct caches not only boolean detection results but also brand scoreboards, CPUID leaf support bits, and hardware concurrency counts. core::run_all()consultsmemo::is_cached()before executing any expensive check, ensuring deterministic operations run once per process.- Subsequent calls to
VM::detect()orVM::brand()retrieve data from read-only caches, eliminating system calls and reducing execution time to microseconds. - The architecture is lock-free after initial population and compatible with early-exit optimization flags.
Frequently Asked Questions
What specific data does VMAware's memoization system cache?
The system caches three categories of data: (1) per-technique boolean results and scoring weights in memo::cache_table, (2) aggregated brand strings and scoreboards used by VM::brand(), and (3) auxiliary hardware information such as CPUID leaf support bits, CPU vendor strings, and thread counts. All caches reside in static storage within [src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3299).
How does memoization affect thread safety?
The memoization layer is lock-free for reads after the first detection pass. The cache_table is populated during the initial single-threaded detection sweep (or under the library's internal mutex if called concurrently), after which all entries are immutable. This design allows multiple threads to call VM::detect() simultaneously without contention, as subsequent accesses only read from the pre-computed static cache.
Can the memoization cache be cleared programmatically?
Yes. While the cache persists for the process lifetime by default, the VM::memo struct provides the uncache(flag) method to invalidate specific technique entries. This is useful when the library is used in long-running processes that need to re-evaluate the environment after major system changes, though the source code in [src/vmaware.hpp](https://github.com/kernelwernel/vmaware/blob/main/src/vmaware.hpp#L3131-L3148) primarily optimizes for the common case of immutable VM detection results.
Does memoization work with VMAware's "short-cut" detection mode?
Absolutely. When the SHORTCUT flag is enabled, detection terminates early once the confidence score exceeds the hypervisor threshold. The memoization system stores results for all techniques executed before the short-cut exit, including partial brand scoreboard data. If VM::detect() is called again, the library instantly returns the cached conclusion without re-scanning, even if the first run stopped early. This combination yields the fastest possible repeated detection path.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →