Performance Implications of Running All 90+ Detection Techniques in VMAware

Running the complete VMAware detection suite with VM::ALL typically incurs 0.5–5 ms of overhead on modern hardware, scaling to approximately 10 ms on low-end systems due to I/O-heavy registry and file-system queries.

VMAware is a cross-platform C++ library for virtual machine detection that ships with over 90 distinct techniques enumerated in VM::enum_flags. When you invoke the exhaustive detection suite using VM::detect(VM::ALL), the library iterates through VM::technique_vector and executes every check sequentially unless the result is already memo-cached. Understanding the performance characteristics of this full scan is critical for applications that require low-latency initialization or repeated VM checks.

How the Full Detection Suite Executes

When you call VM::detect(VM::ALL), the library triggers core::run_all() in src/vmaware.hpp, which walks through the entire technique vector. Each technique is executed once, with results stored in an internal memoisation cache (memo::leaf_cache, memo::cpu_brand, etc.) to ensure subsequent calls are O(1) lookups.

The execution flow follows this pattern:

  1. The library checks the memo cache for existing results
  2. For cache misses, it executes the technique-specific detection logic
  3. CPU-only checks read registers directly
  4. I/O-heavy checks query /sys/devices/virtual/dmi/id/*, Windows registry hives, or spawn external processes
  5. Results are cached for future invocations

Because VMAware runs techniques sequentially in a single thread, the total latency is the sum of all individual technique costs. There is no built-in parallel execution, so the full scan duration depends entirely on the cumulative cost of the selected detection methods.

Performance Breakdown by Technique Category

The 90+ techniques fall into distinct performance tiers based on their system interaction requirements:

CPU-Only Checks (e.g., VM::CPU_BRAND, VM::HYPERVISOR_BIT, RDTSC timing)

  • Simple register reads and arithmetic operations
  • Cost: Tens to hundreds of nanoseconds per technique
  • These dominate the technique vector but contribute minimally to total runtime

File-System and Registry Queries (e.g., reading /proc, /sys files, Windows BIOS tables, systeminfo execution)

  • Require open-read-close cycles or registry hive access
  • Cost: Few microseconds to ~1 ms per technique, depending on I/O latency and cache state

Network-Oriented or External-Process Checks (e.g., WMI queries, MAC address prefix probing, disk serial enumeration via VM::DISK_SERIAL)

  • System calls, process spawns, and privilege escalation for hardware access
  • Cost: ~1 ms to 10 ms for the slowest checks

Memoisation Impact

  • First invocation pays the full I/O cost
  • Subsequent calls retrieve cached results from memo:: structures
  • Cost: Effectively zero for repeated checks regardless of original technique cost

Benchmark Data from auxiliary/benchmark.cpp

The repository includes a dedicated benchmark utility at auxiliary/benchmark.cpp that measures individual technique latency using high-resolution timers. Typical output on contemporary hardware shows:

// Example benchmark output (first run, cache cold)
VM::CPU_BRAND:      0.12 µs
VM::HYPERVISOR_BIT: 0.08 µs
VM::DISK_SERIAL:    1.47 ms
VM::detect():       0.68 ms  (VM::DEFAULT)
VM::detect(ALL):    2.34 ms  (VM::ALL)

Typical one-off detection using VM::ALL completes in 0.5–5 ms on modern PCs. On low-end hardware or systems with high I/O latency, this can extend to ~10 ms. The benchmark demonstrates that while most techniques are microsecond-scale, a few I/O-heavy checks dominate the total runtime.

To run the benchmark yourself:


# Compile from the repository root

g++ -std=c++17 -O2 auxiliary/benchmark.cpp -o benchmark
./benchmark

False-Positive Risks with VM::ALL

Enabling all 90+ techniques activates several checks that are intentionally disabled by default in VM::DEFAULT. These techniques are excluded from the default set because they produce false positives on certain hypervisors or sandbox environments. While this does not directly impact timing performance, it affects detection accuracy:

  • VM::ALL includes noisy techniques that may fire on bare-metal systems with specific hardware configurations
  • The default flag set (VM::DEFAULT) provides a balanced 95%+ detection rate with minimal false positives
  • Consider the accuracy trade-off when selecting the exhaustive suite for security-sensitive applications

Optimization Strategies for Production

For applications where latency matters, avoid the brute-force approach:

Use VM::DEFAULT for Balanced Detection

// Fast, well-balanced set (~0.5 ms)
bool is_vm = VM::detect();  // Equivalent to VM::detect(VM::DEFAULT)

Selectively Pick Cheap Techniques

// Explicitly select CPU-only checks for microsecond-scale detection
bool quick_check = VM::detect(VM::CPU_BRAND | VM::HYPERVISOR_BIT);

Leverage Memoisation for Repeated Checks

// First call pays full cost
bool first = VM::detect(VM::ALL);

// Subsequent calls are O(1) lookups from memo:: cache
bool second = VM::detect(VM::ALL);  // Effectively free

Avoid VM::ALL in Performance-Critical Loops

  • Use VM::DEFAULT or explicit technique flags for repeated polling
  • Reserve VM::ALL for one-off initialization checks where a few milliseconds of overhead is acceptable

Summary

  • Full detection cost: Running VM::detect(VM::ALL) executes all 90+ techniques sequentially, costing 0.5–10 ms depending on hardware and I/O latency.
  • I/O dominates: File-system queries (DISK_SERIAL, registry checks) contribute the most latency, while CPU register reads complete in nanoseconds.
  • Memoisation mitigates: The memo:: caching layer in src/vmaware.hpp ensures repeated calls are free after the first execution.
  • Accuracy trade-off: VM::ALL includes noisy techniques disabled by default, increasing false-positive risk for marginal detection gains.
  • Production recommendation: Use VM::DEFAULT for balanced performance, or explicitly select cheap techniques for sub-millisecond detection.

Frequently Asked Questions

How long does VM::detect(VM::ALL) take to execute?

According to the benchmark utility in auxiliary/benchmark.cpp, VM::detect(VM::ALL) typically completes in 0.5–5 ms on modern hardware, with the potential to reach ~10 ms on low-end systems or when disk I/O is slow. The first invocation incurs the full cost, while subsequent calls retrieve memo-cached results in nanoseconds.

Does VMAware parallelize the 90+ detection techniques?

No. As implemented in core::run_all() within src/vmaware.hpp, VMAware executes techniques sequentially in a single thread. The total runtime equals the sum of individual technique costs. There is currently no built-in multi-threading or parallel execution for the detection suite.

Why is VM::ALL slower than the default detection?

VM::ALL activates file-system and registry queries that VM::DEFAULT excludes. Techniques like VM::DISK_SERIAL and WMI queries require spawning processes or reading BIOS tables, which cost 1–10 ms compared to <1 µs for CPU-only checks. The default set deliberately omits these I/O-heavy techniques to maintain sub-millisecond performance.

Can I run specific techniques to improve performance?

Yes. Instead of VM::ALL, pass a bitwise OR of specific flags to VM::detect(). For example, VM::detect(VM::CPU_BRAND | VM::HYPERVISOR_BIT) executes only register-based checks, completing in <1 µs. Refer to VM::enum_flags in src/vmaware.hpp for the full list of available technique flags.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →