# How Ghidra's P-Code Intermediate Language Works: A Deep Dive into the SLEIGH Translation Layer

> Uncover Ghidra's P-Code intermediate language. Learn how SLEIGH translates binary instructions into uniform micro-operations for powerful analysis and emulation across architectures.

- Repository: [National Security Agency/ghidra](https://github.com/NationalSecurityAgency/ghidra)
- Tags: deep-dive
- Published: 2026-03-04

---

**Ghidra's P-Code is a processor-independent intermediate representation that translates native binary instructions into three-address micro-operations, enabling uniform analysis and emulation across all supported architectures.**

Ghidra, the open-source reverse engineering framework developed by the National Security Agency, uses P-Code as its foundational intermediate language to abstract away processor-specific details. This intermediate representation allows analysts to write architecture-agnostic analysis tools and emulators. Understanding how Ghidra's P-Code intermediate language works requires examining the translation pipeline from SLEIGH specifications through to executable `PcodeOp` sequences.

## What Is Ghidra's P-Code Intermediate Language?

P-Code is a **register transfer language (RTL)** that represents every native processor instruction as a sequence of simple, three-address operations. Unlike assembly, which varies by architecture, P-Code provides a consistent vocabulary of approximately 60 operations (defined in [`PcodeOp.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeOp.java)) that cover arithmetic, logic, memory access, and control flow. Each operation manipulates **Varnodes**—typed storage locations that abstract registers, memory addresses, and constants.

## From Native Instructions to P-Code: The SLEIGH Translation Pipeline

The translation from binary to P-Code begins with **SLEIGH**, Ghidra's processor specification language. When Ghidra disassembles an instruction, the `SleighBase` class invokes `PcodeTranslate` to emit a list of `PcodeOp` objects.

### Parsing and Symbol Resolution

At the heart of this process is [`PcodeParser.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeParser.java), which builds an intermediate representation using `VectorSTL<OpTpl>` trees. The parser resolves symbolic references—including special symbols like `inst_start` and `inst_next`—before creating concrete `PcodeOp` instances:

```java
// In PcodeParser.java – building the symbol table used during translation
symbolMap.put("inst_start", new StartSymbol(internalLoc, "inst_start", getConstantSpace()));
symbolMap.put("inst_next",  new EndSymbol(internalLoc, "inst_next", getConstantSpace()));

```

These symbols allow p-code operations to reference the current instruction address and the fall-through address, essential for relative addressing and control flow analysis.

## The P-Code Data Model: PcodeOp and Varnode

Every p-code operation is encapsulated in a **`PcodeOp`** object (defined in [`PcodeOp.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeOp.java), lines 45-138), which contains four critical fields:

- **`opcode`**: An integer constant representing the operation (e.g., `COPY = 1`, `INT_ADD = 19`, `CALL = 7`)
- **`input[]`**: An array of `Varnode` objects representing source operands
- **`output`**: A destination `Varnode` (or `null` for operations like `CALL` that don't produce results)
- **`seqnum`**: A `SequenceNumber` that uniquely identifies the operation within a function's execution context

A **`Varnode`** (defined in [`Varnode.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/Varnode.java)) abstracts storage locations through three attributes: size in bytes, address space (register, RAM, or constant), and offset within that space.

### Programmatic Access with PcodeProgram

The `PcodeProgram` class (lines 21-45 in [`PcodeProgram.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeProgram.java)) serves as a container for a sequence of `PcodeOp` objects along with a map of user-defined operations. You can generate a program directly from a Ghidra instruction:

```java
// Create a PcodeProgram from a concrete instruction
PcodeProgram prog = PcodeProgram.fromInstruction(instr);

```

## Executing P-Code: The PcodeExecutor Kernel

Emulation and analysis of p-code programs are handled by **`PcodeExecutor`** (implemented in [`PcodeExecutor.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeExecutor.java), lines 34-80). This kernel implements a fetch-decode-execute cycle that processes each `PcodeOp` sequentially through two pluggable interfaces:

- **`PcodeArithmetic<T>`**: Implements arithmetic and logical operations (like `INT_ADD`, `BOOL_AND`) for a specific value type `T`
- **`PcodeExecutorState<T>`**: Handles memory and register access for `LOAD`, `STORE`, and temporary variable operations

The executor creates a `PcodeFrame` to track the program counter and repeatedly calls `step()` until the frame's `finish()` method signals completion. A typical setup uses the default byte-array implementations:

```java
PcodeExecutorState<byte[]> state = new DefaultPcodeExecutorState<>(addressFactory);
PcodeArithmetic<byte[]> arithmetic = new DefaultPcodeArithmetic<>(addressFactory);
PcodeExecutor<byte[]> executor =
        new PcodeExecutor<>(sleighLanguage, arithmetic, state, Reason.USEROP);
executor.execute(prog, PcodeUseropLibrary.nil());

```

## Extending P-Code with User Operations (UserOps)

Ghidra allows processor specifications to define custom operations beyond the standard p-code instruction set through **`UserOpSymbol`** and the **`PcodeUseropLibrary`** interface (defined in [`PcodeUseropLibrary.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeUseropLibrary.java)). When the executor encounters a `CALLOTHER` opcode, it resolves the user-op index through the library.

The default implementation (`PcodeUseropLibrary.nil()`) throws an exception for undefined user-ops, but analysts can inject custom Java logic:

```java
public class MyUserOps implements PcodeUseropDefinition<byte[]> {
    @Override
    public String getName() { return "my_print"; }

    @Override
    public int getNumberOfInputs() { return 1; }

    @Override
    public int getNumberOfOutputs() { return 0; }

    @Override
    public void evaluate(PcodeExecutorState<byte[]> state,
                         List<byte[]> inputs) {
        System.out.println("Custom value: " + java.util.Arrays.toString(inputs.get(0)));
    }
}

// Register the op (assume language defines a CALLOTHER with index 0)
PcodeUseropLibrary<byte[]> lib = new DefaultPcodeUseropLibrary<>();
lib.registerUserop(0, new MyUserOps());
exec.execute(program, lib);

```

This mechanism enables emulation of processor-specific instructions like vector operations or system calls by implementing the behavior in Java rather than SLEIGH.

## Practical Examples: Working with Ghidra's P-Code Intermediate Language

### Printing P-Code for an Instruction

To analyze the p-code generated for a specific instruction:

```java
import ghidra.program.model.listing.*;
import ghidra.program.model.pcode.*;

public void dumpInstructionPcode(Instruction instr) {
    // Build a PcodeProgram (no user-ops)
    PcodeProgram prog = PcodeProgram.fromInstruction(instr);
    System.out.println("P-code for " + instr + ":");
    int idx = 0;
    for (PcodeOp op : prog.code) {
        System.out.printf("%02d: %-15s %s\n",
            idx++,
            PcodeOp.getMnemonic(op.getOpcode()),
            op);
    }
}

```

### Emulating Register Addition

For low-level emulation, you can construct `PcodeOp` objects manually and execute them:

```java
import ghidra.pcode.exec.*;
import ghidra.pcode.exec.PcodeExecutorState.*;
import ghidra.pcode.emu.*;

public byte[] emulateAdd(Register r1, Register r2, Register out,
                        SleighLanguage lang, AddressFactory af) {
    // 1️⃣ Build a tiny p-code program manually
    Varnode in1 = new Varnode(r1.getAddress(), r1.getMinimumByteSize());
    Varnode in2 = new Varnode(r2.getAddress(), r2.getMinimumByteSize());
    Varnode dst = new Varnode(out.getAddress(), out.getMinimumByteSize());

    PcodeOp add = new PcodeOp(new SequenceNumber(out.getAddress(), 0),
                              PcodeOp.INT_ADD, new Varnode[]{in1, in2}, dst);
    PcodeProgram prog = new PcodeProgram(lang, List.of(add), Map.of());

    // 2️⃣ Provide an executor state with concrete byte[] values
    DefaultPcodeExecutorState<byte[]> state = new DefaultPcodeExecutorState<>(af);
    state.setVar(in1, new byte[]{0x05});   // r1 = 5
    state.setVar(in2, new byte[]{0x03});   // r2 = 3

    // 3️⃣ Execute
    PcodeExecutor<byte[]> exec = new PcodeExecutor<>(lang,
            new DefaultPcodeArithmetic<>(af), state, Reason.USEROP);
    exec.execute(prog, PcodeUseropLibrary.nil());

    // 4️⃣ Retrieve result
    return state.getVar(dst);
}

```

## Summary

- **P-Code** provides a processor-independent intermediate language that unifies analysis across Ghidra's supported architectures
- **SLEIGH specifications** drive the translation from binary instructions to `PcodeOp` sequences via [`PcodeParser.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeParser.java)
- The **PcodeOp** model uses **Varnodes** to represent typed storage locations with explicit opcodes, inputs, outputs, and sequence numbers
- **PcodeExecutor** implements a plug-in architecture separating arithmetic logic (`PcodeArithmetic`) from state management (`PcodeExecutorState`)
- **UserOps** extend the instruction set through `PcodeUseropLibrary`, allowing custom Java implementations of processor-specific operations

## Frequently Asked Questions

### What is the difference between P-Code and LLVM IR?

While both are intermediate representations, **P-Code** is specifically designed for binary analysis and reverse engineering, operating at a lower abstraction level closer to hardware. LLVM IR targets compiler optimization with high-level type information and SSA form, whereas P-Code (as implemented in [`PcodeOp.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeOp.java)) focuses on explicit memory operations and register transfers without assuming high-level language semantics.

### How does Ghidra's P-Code handle floating-point operations?

Floating-point operations are supported through specific opcodes like `FLOAT_ADD` and `FLOAT_MULT` defined in the `PcodeOp` constants. The `PcodeArithmetic<T>` interface implementations (such as `DefaultPcodeArithmetic`) provide the actual floating-point semantics, typically delegating to Java's floating-point operations when using byte-array value types.

### Can I modify P-Code after it has been generated from an instruction?

Yes, analysts can construct modified `PcodeProgram` instances by manipulating the `List<PcodeOp>` before execution. While you cannot alter the p-code cached within Ghidra's database for existing instructions, you can create new `PcodeOp` instances and assemble them into programs for emulation or analysis purposes using the `PcodeProgram` constructor.

### Where are the P-Code opcodes defined in the Ghidra source code?

The complete opcode enumeration is defined in **[`PcodeOp.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/PcodeOp.java)** at [`Ghidra/Framework/SoftwareModeling/src/main/java/ghidra/program/model/pcode/PcodeOp.java`](https://github.com/NationalSecurityAgency/ghidra/blob/main/Ghidra/Framework/SoftwareModeling/src/main/java/ghidra/program/model/pcode/PcodeOp.java) (lines 45-138). This file contains integer constants for all operations ranging from `COPY` and `LOAD` to specialized operations like `CALLOTHER` for user-defined ops.