How to Use Capstone for Disassembly in iOS VM Patching: A Complete Workflow
The vPhone-CLI project leverages the Capstone disassembly framework to analyze ARM64 Mach-O binaries inside iOS virtual machines, enabling semantic pattern recognition and precise binary patching through detailed operand inspection.
The vPhone-CLI repository provides open-source tools for modifying iOS kernel and userspace binaries within a virtual iPhone environment. To locate code-signing checks and security-critical instruction idioms, developers use Capstone—a fast, multi-architecture disassembler—to decode ARM64 machine code with full semantic context rather than relying on brittle hard-coded offsets.
Initializing the Capstone Engine for ARM64
The patching pipeline begins by creating a global Capstone instance configured for Apple Silicon. In scripts/patchers/cfw_asm.py, the initialization code sets the architecture to ARM64 with little-endian mode and enables detailed operand decoding.
from capstone import Cs, CS_ARCH_ARM64, CS_MODE_LITTLE_ENDIAN
# Global engine initialization as implemented in cfw_asm.py
_cs = Cs(CS_ARCH_ARM64, CS_MODE_LITTLE_ENDIAN)
_cs.detail = True # Enable ARM64_OP_REG, ARM64_OP_IMM access
Enabling _cs.detail = True is critical for iOS VM patching. It allows subsequent analysis to distinguish between register operands (ARM64_OP_REG) and immediate values (ARM64_OP_IMM), which is necessary for identifying specific registers like w0 or x0 in security checks.
Loading and Addressing Mach-O Binaries
Before disassembly, the target binary must be mapped into a mutable buffer. The DSCChunks helper in scripts/patchers/cfw_dsc_chunks.py provides virtual-memory-aware access via bytes_at_vma() and write_at_vma() methods.
The patchers load the Mach-O file into a bytearray, allowing in-place modification once Capstone identifies the patch sites. This approach maintains the alignment between file offsets and virtual addresses (VMA), ensuring the disassembler interprets instruction boundaries correctly relative to the iOS runtime layout.
Disassembling Byte Ranges with disasm_at()
The core disassembly logic resides in the disasm_at() function within scripts/patchers/cfw_asm.py. This helper slices the binary at a specific offset and returns a list of CsInsn objects containing address, mnemonic, op_str, and the detailed operands list.
# From cfw_asm.py (lines 84-86)
def disasm_at(data, off, n=8):
"""Disassemble n instructions starting at file offset off."""
return list(_cs.disasm(bytes(data[off:off+n*4]), off))
Each CsInsn object provides the instruction's virtual address and mnemonic, while the operands array exposes register names via insn.reg_name(). This granularity allows patchers to verify that a cset instruction targets a specific condition flag or that an eor operation uses the expected source registers.
Semantic Pattern Matching via Operand Inspection
Rather than scanning for byte signatures that break across compiler versions, vPhone-CLI uses semantic pattern matching. The helper function _find_consistency_check() in scripts/patchers/cfw_patch_xpc_lwcr.py (lines 109-148) walks the instruction list to locate the cset-eor-tbz idiom used in XPC lightweight code requirement (LWCR) validation.
# Pattern matching logic derived from cfw_patch_xpc_lwcr.py
def find_lwcr_idiom(insns):
"""Locate the cset wC, ne; eor wE, w0, wC; tbz wE, #0 sequence."""
for i in range(len(insns) - 1):
eor = insns[i]
if eor.mnemonic != "eor":
continue
# Extract register names using _rn() helper
regs = [_rn(eor, j) for j in range(3)]
if regs[1] != "w0":
continue
tbz = insns[i + 1]
if tbz.mnemonic != "tbz" or _rn(tbz, 0) != regs[0]:
continue
# Walk backwards to find matching cset
for j in range(i - 1, max(-1, i - 6), -1):
cand = insns[j]
if cand.mnemonic == "cset" and _rn(cand, 0) == regs[2]:
if cand.op_str.strip().endswith("ne"):
return cand, eor, tbz
return None
By inspecting operand types and register names rather than raw bytes, the patcher remains robust against register reallocation or surrounding instruction changes introduced by newer iOS SDKs.
Patching with Keystone Integration
Once Capstone identifies the target instructions, the patcher generates replacement bytes using Keystone (the assembling counterpart to Capstone). The asm() function in cfw_asm.py assembles ARM64 instructions like nop or cset w0, eq into machine code bytes.
The patch_xpc_lwcr() function (lines 84-92 in cfw_patch_xpc_lwcr.py) builds an edits list mapping virtual addresses to new byte sequences, then writes these back to the Mach-O via DSCChunks.write_at_vma(). This workflow ensures deterministic, byte-for-byte replacements without external symbol maps or debug information.
Summary
- Architecture-aware decoding: Capstone's ARM64 backend with
detail=Trueexposes register and immediate operands, enabling semantic analysis of security-critical code paths. - Virtual-memory-aligned disassembly: The
disasm_at()helper incfw_asm.pyoperates on file slices mapped to virtual addresses, reproducing the iOS runtime instruction layout. - Semantic pattern matching: Functions like
_find_consistency_check()identify instruction idioms (e.g.,cset-eor-tbz) by inspecting operand types rather than hard-coded offsets. - Keystone integration: Discovered patterns are replaced with assembled instructions via
DSCChunks, creating upgrade-resistant patches for kernel and userspace binaries.
Frequently Asked Questions
What makes Capstone suitable for iOS binary patching?
Capstone provides deterministic, lightweight disassembly for ARM64 with detailed operand information, allowing the vPhone-CLI tools to distinguish between registers and immediates. This precision is necessary for identifying security checks that use specific registers like w0 or x0, which simpler byte-pattern scanners cannot reliably isolate.
How does vPhone-CLI handle Mach-O virtual addresses during disassembly?
The DSCChunks class in cfw_dsc_chunks.py maintains a mapping between file offsets and virtual memory addresses (VMA). When disasm_at() processes a range, it uses the VMA as the base address for Capstone, ensuring that branch targets and symbol references align with the iOS VM's runtime layout.
Why use semantic pattern matching instead of byte signatures?
Byte signatures break when Apple changes compiler optimizations or register allocation in iOS updates. By analyzing instruction semantics—such as verifying that an eor instruction uses w0 as the second operand—the patchers remain functional across OS versions without manual signature updates.
Can this workflow patch ARM64e (pointer authentication) binaries?
Yes, the Capstone-based approach works on ARM64e binaries because it operates at the instruction level before pointer authentication codes (PAC) are applied to pointers. However, patches must respect ARM64e ABI constraints, and the Keystone assembler must generate proper authenticated instructions when patching function prologues or pointers specifically protected by PAC.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →