How Ghidra's P-Code Intermediate Language Works: A Deep Dive into the SLEIGH Translation Layer

Ghidra's P-Code is a processor-independent intermediate representation that translates native binary instructions into three-address micro-operations, enabling uniform analysis and emulation across all supported architectures.

Ghidra, the open-source reverse engineering framework developed by the National Security Agency, uses P-Code as its foundational intermediate language to abstract away processor-specific details. This intermediate representation allows analysts to write architecture-agnostic analysis tools and emulators. Understanding how Ghidra's P-Code intermediate language works requires examining the translation pipeline from SLEIGH specifications through to executable PcodeOp sequences.

What Is Ghidra's P-Code Intermediate Language?

P-Code is a register transfer language (RTL) that represents every native processor instruction as a sequence of simple, three-address operations. Unlike assembly, which varies by architecture, P-Code provides a consistent vocabulary of approximately 60 operations (defined in PcodeOp.java) that cover arithmetic, logic, memory access, and control flow. Each operation manipulates Varnodes—typed storage locations that abstract registers, memory addresses, and constants.

From Native Instructions to P-Code: The SLEIGH Translation Pipeline

The translation from binary to P-Code begins with SLEIGH, Ghidra's processor specification language. When Ghidra disassembles an instruction, the SleighBase class invokes PcodeTranslate to emit a list of PcodeOp objects.

Parsing and Symbol Resolution

At the heart of this process is PcodeParser.java, which builds an intermediate representation using VectorSTL<OpTpl> trees. The parser resolves symbolic references—including special symbols like inst_start and inst_next—before creating concrete PcodeOp instances:

// In PcodeParser.java – building the symbol table used during translation
symbolMap.put("inst_start", new StartSymbol(internalLoc, "inst_start", getConstantSpace()));
symbolMap.put("inst_next",  new EndSymbol(internalLoc, "inst_next", getConstantSpace()));

These symbols allow p-code operations to reference the current instruction address and the fall-through address, essential for relative addressing and control flow analysis.

The P-Code Data Model: PcodeOp and Varnode

Every p-code operation is encapsulated in a PcodeOp object (defined in PcodeOp.java, lines 45-138), which contains four critical fields:

  • opcode: An integer constant representing the operation (e.g., COPY = 1, INT_ADD = 19, CALL = 7)
  • input[]: An array of Varnode objects representing source operands
  • output: A destination Varnode (or null for operations like CALL that don't produce results)
  • seqnum: A SequenceNumber that uniquely identifies the operation within a function's execution context

A Varnode (defined in Varnode.java) abstracts storage locations through three attributes: size in bytes, address space (register, RAM, or constant), and offset within that space.

Programmatic Access with PcodeProgram

The PcodeProgram class (lines 21-45 in PcodeProgram.java) serves as a container for a sequence of PcodeOp objects along with a map of user-defined operations. You can generate a program directly from a Ghidra instruction:

// Create a PcodeProgram from a concrete instruction
PcodeProgram prog = PcodeProgram.fromInstruction(instr);

Executing P-Code: The PcodeExecutor Kernel

Emulation and analysis of p-code programs are handled by PcodeExecutor (implemented in PcodeExecutor.java, lines 34-80). This kernel implements a fetch-decode-execute cycle that processes each PcodeOp sequentially through two pluggable interfaces:

  • PcodeArithmetic<T>: Implements arithmetic and logical operations (like INT_ADD, BOOL_AND) for a specific value type T
  • PcodeExecutorState<T>: Handles memory and register access for LOAD, STORE, and temporary variable operations

The executor creates a PcodeFrame to track the program counter and repeatedly calls step() until the frame's finish() method signals completion. A typical setup uses the default byte-array implementations:

PcodeExecutorState<byte[]> state = new DefaultPcodeExecutorState<>(addressFactory);
PcodeArithmetic<byte[]> arithmetic = new DefaultPcodeArithmetic<>(addressFactory);
PcodeExecutor<byte[]> executor =
        new PcodeExecutor<>(sleighLanguage, arithmetic, state, Reason.USEROP);
executor.execute(prog, PcodeUseropLibrary.nil());

Extending P-Code with User Operations (UserOps)

Ghidra allows processor specifications to define custom operations beyond the standard p-code instruction set through UserOpSymbol and the PcodeUseropLibrary interface (defined in PcodeUseropLibrary.java). When the executor encounters a CALLOTHER opcode, it resolves the user-op index through the library.

The default implementation (PcodeUseropLibrary.nil()) throws an exception for undefined user-ops, but analysts can inject custom Java logic:

public class MyUserOps implements PcodeUseropDefinition<byte[]> {
    @Override
    public String getName() { return "my_print"; }

    @Override
    public int getNumberOfInputs() { return 1; }

    @Override
    public int getNumberOfOutputs() { return 0; }

    @Override
    public void evaluate(PcodeExecutorState<byte[]> state,
                         List<byte[]> inputs) {
        System.out.println("Custom value: " + java.util.Arrays.toString(inputs.get(0)));
    }
}

// Register the op (assume language defines a CALLOTHER with index 0)
PcodeUseropLibrary<byte[]> lib = new DefaultPcodeUseropLibrary<>();
lib.registerUserop(0, new MyUserOps());
exec.execute(program, lib);

This mechanism enables emulation of processor-specific instructions like vector operations or system calls by implementing the behavior in Java rather than SLEIGH.

Practical Examples: Working with Ghidra's P-Code Intermediate Language

Printing P-Code for an Instruction

To analyze the p-code generated for a specific instruction:

import ghidra.program.model.listing.*;
import ghidra.program.model.pcode.*;

public void dumpInstructionPcode(Instruction instr) {
    // Build a PcodeProgram (no user-ops)
    PcodeProgram prog = PcodeProgram.fromInstruction(instr);
    System.out.println("P-code for " + instr + ":");
    int idx = 0;
    for (PcodeOp op : prog.code) {
        System.out.printf("%02d: %-15s %s\n",
            idx++,
            PcodeOp.getMnemonic(op.getOpcode()),
            op);
    }
}

Emulating Register Addition

For low-level emulation, you can construct PcodeOp objects manually and execute them:

import ghidra.pcode.exec.*;
import ghidra.pcode.exec.PcodeExecutorState.*;
import ghidra.pcode.emu.*;

public byte[] emulateAdd(Register r1, Register r2, Register out,
                        SleighLanguage lang, AddressFactory af) {
    // 1️⃣ Build a tiny p-code program manually
    Varnode in1 = new Varnode(r1.getAddress(), r1.getMinimumByteSize());
    Varnode in2 = new Varnode(r2.getAddress(), r2.getMinimumByteSize());
    Varnode dst = new Varnode(out.getAddress(), out.getMinimumByteSize());

    PcodeOp add = new PcodeOp(new SequenceNumber(out.getAddress(), 0),
                              PcodeOp.INT_ADD, new Varnode[]{in1, in2}, dst);
    PcodeProgram prog = new PcodeProgram(lang, List.of(add), Map.of());

    // 2️⃣ Provide an executor state with concrete byte[] values
    DefaultPcodeExecutorState<byte[]> state = new DefaultPcodeExecutorState<>(af);
    state.setVar(in1, new byte[]{0x05});   // r1 = 5
    state.setVar(in2, new byte[]{0x03});   // r2 = 3

    // 3️⃣ Execute
    PcodeExecutor<byte[]> exec = new PcodeExecutor<>(lang,
            new DefaultPcodeArithmetic<>(af), state, Reason.USEROP);
    exec.execute(prog, PcodeUseropLibrary.nil());

    // 4️⃣ Retrieve result
    return state.getVar(dst);
}

Summary

  • P-Code provides a processor-independent intermediate language that unifies analysis across Ghidra's supported architectures
  • SLEIGH specifications drive the translation from binary instructions to PcodeOp sequences via PcodeParser.java
  • The PcodeOp model uses Varnodes to represent typed storage locations with explicit opcodes, inputs, outputs, and sequence numbers
  • PcodeExecutor implements a plug-in architecture separating arithmetic logic (PcodeArithmetic) from state management (PcodeExecutorState)
  • UserOps extend the instruction set through PcodeUseropLibrary, allowing custom Java implementations of processor-specific operations

Frequently Asked Questions

What is the difference between P-Code and LLVM IR?

While both are intermediate representations, P-Code is specifically designed for binary analysis and reverse engineering, operating at a lower abstraction level closer to hardware. LLVM IR targets compiler optimization with high-level type information and SSA form, whereas P-Code (as implemented in PcodeOp.java) focuses on explicit memory operations and register transfers without assuming high-level language semantics.

How does Ghidra's P-Code handle floating-point operations?

Floating-point operations are supported through specific opcodes like FLOAT_ADD and FLOAT_MULT defined in the PcodeOp constants. The PcodeArithmetic<T> interface implementations (such as DefaultPcodeArithmetic) provide the actual floating-point semantics, typically delegating to Java's floating-point operations when using byte-array value types.

Can I modify P-Code after it has been generated from an instruction?

Yes, analysts can construct modified PcodeProgram instances by manipulating the List<PcodeOp> before execution. While you cannot alter the p-code cached within Ghidra's database for existing instructions, you can create new PcodeOp instances and assemble them into programs for emulation or analysis purposes using the PcodeProgram constructor.

Where are the P-Code opcodes defined in the Ghidra source code?

The complete opcode enumeration is defined in PcodeOp.java at Ghidra/Framework/SoftwareModeling/src/main/java/ghidra/program/model/pcode/PcodeOp.java (lines 45-138). This file contains integer constants for all operations ranging from COPY and LOAD to specialized operations like CALLOTHER for user-defined ops.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →