# What Is LLVM IR and How to Optimize It Programmatically

> Understand LLVM IR, the compiler infrastructure's intermediate representation. Learn to programmatically optimize LLVM IR using the PassBuilder API to create custom C++ transformation pipelines.

- Repository: [LLVM/llvm-project](https://github.com/llvm/llvm-project)
- Tags: deep-dive
- Published: 2026-09-11

---

**LLVM IR is a language-agnostic, static single assignment (SSA) based intermediate representation that serves as the backbone of the LLVM compiler infrastructure, and you can optimize it programmatically using the `PassBuilder` API to construct custom transformation pipelines in C++.**

LLVM IR (Intermediate Representation) is the core data format that powers the entire LLVM ecosystem. Stored in the `llvm/llvm-project` repository, this SSA-based representation enables front-ends like Clang to generate portable, type-safe code that can be aggressively optimized before target-specific code generation.

## Core Architecture of LLVM IR

LLVM IR provides a universal abstraction layer between source languages and target architectures.

### Modules, Functions, and Basic Blocks

The top-level container is the **`llvm::Module`**, which owns a collection of **`llvm::Function`** objects. Each function consists of a list of **`llvm::BasicBlock`** instances, forming a control-flow graph. According to the source in [`llvm/include/llvm/IR/Module.h`](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/Module.h), the module maintains a symbol table for global variables and function declarations.

### SSA Form and Type System

Every instruction in LLVM IR produces a typed **`llvm::Value`** in strict SSA form, meaning each variable is defined exactly once. The **`llvm::IRBuilder`** class—defined in [[`llvm/include/llvm/IR/IRBuilder.h`](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/IRBuilder.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/IRBuilder.h)—automates the creation of these values while managing insertion points and type inference.

## Generating LLVM IR Programmatically

You can produce LLVM IR either through compiler front-ends or direct API construction.

### Emitting IR with Clang

The simplest method uses Clang to lower C/C++ into human-readable IR:

```bash
clang -S -emit-llvm -o output.ll input.c

```

This generates a `.ll` file containing the textual representation of the intermediate code.

### Building IR with IRBuilder

For custom code generation, the `llvm::IRBuilder` template class provides a fluent API. As implemented in [[`llvm/include/llvm/IR/IRBuilder.h`](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/IRBuilder.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/IRBuilder.h), this utility handles instruction insertion, automatic naming, and debug location tracking.

The following example constructs a simple addition function:

```cpp
#include "llvm/IR/IRBuilder.h"
#include "llvm/IR/Verifier.h"
#include "llvm/Support/raw_ostream.h"

int main() {
  llvm::LLVMContext Ctx;
  llvm::Module Mod("demo", Ctx);
  llvm::IRBuilder<> Builder(Ctx);

  // Define function type: int (int, int)
  auto *Int32Ty = Builder.getInt32Ty();
  auto *FuncTy = llvm::FunctionType::get(Int32Ty, {Int32Ty, Int32Ty}, false);
  auto *Func = llvm::Function::Create(FuncTy, 
                                      llvm::Function::ExternalLinkage,
                                      "add", Mod);

  // Name the arguments
  auto ArgIt = Func->arg_begin();
  ArgIt->setName("a");
  (++ArgIt)->setName("b");

  // Create entry block
  llvm::BasicBlock *Entry = llvm::BasicBlock::Create(Ctx, "entry", Func);
  Builder.SetInsertPoint(Entry);

  // Generate add instruction: %sum = add i32 %a, %b
  llvm::Value *Sum = Builder.CreateAdd(Func->getArg(0), 
                                       Func->getArg(1), "sum");
  Builder.CreateRet(Sum);

  // Verify and print
  llvm::verifyFunction(*Func);
  Mod.print(llvm::outs(), nullptr);
}

```

## Programmatic Optimization Pipelines

Once you have generated IR, you can optimize it without invoking external tools by using LLVM’s modern pass management infrastructure.

### The PassBuilder Infrastructure

The **`llvm::PassBuilder`** class—declared in [[`llvm/include/llvm/IR/PassManager.h`](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/PassManager.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/PassManager.h)—constructs optimization pipelines for both function-level and module-level transformations. Unlike the legacy pass manager, this system uses explicit analysis managers and preserved analysis tokens to minimize recomputation.

### Constructing and Running Passes

The typical workflow involves instantiating a `PassBuilder`, registering analysis passes, and building pipelines that run transformations like `InstCombinePass` (instruction combining). The `InstCombinePass` header is located at [[`llvm/include/llvm/Transforms/Scalar/InstCombine.h`](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/Transforms/Scalar/InstCombine.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/Transforms/Scalar/InstCombine.h) and performs peephole optimizations on algebraic identities.

Here is a complete example of running an optimization pipeline:

```cpp
#include "llvm/IR/PassManager.h"
#include "llvm/IR/LegacyPassManagers.h"
#include "llvm/Transforms/Scalar/InstCombine.h"

int main() {
  // Assume 'Mod' is a valid llvm::Module from previous construction
  
  llvm::PassBuilder PB(Mod.getContext());

  // Build function-level pipeline
  llvm::FunctionPassManager<> FPM = PB.buildFunctionPassPipeline();
  FPM.addPass(llvm::InstCombinePass());

  // Run on each function
  for (auto &F : Mod) {
    FPM.run(F, llvm::FunctionAnalysisManager());
  }

  // Build and run module-level pipeline
  llvm::ModulePassManager MPM = PB.buildModulePassPipeline();
  MPM.run(Mod, llvm::ModuleAnalysisManager());
}

```

This approach allows you to **compose** specific optimizations—such as aggressive instruction combining, global variable optimization, and control-flow simplification—into a single programmatic execution.

## Summary

- **LLVM IR** is a statically-typed, SSA-based intermediate representation stored in `llvm::Module` objects composed of functions and basic blocks.
- The **`llvm::IRBuilder`** API in [[`IRBuilder.h`](https://github.com/llvm/llvm-project/blob/main/IRBuilder.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/IRBuilder.h) provides the primary interface for programmatically constructing instructions and values.
- The **`llvm::PassBuilder`** class in [[`PassManager.h`](https://github.com/llvm/llvm-project/blob/main/PassManager.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/PassManager.h) manages the modern optimization pipeline infrastructure.
- You can execute transformations like **`InstCombinePass`** from [[`InstCombine.h`](https://github.com/llvm/llvm-project/blob/main/InstCombine.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/Transforms/Scalar/InstCombine.h) directly through the C++ API without spawning external processes.

## Frequently Asked Questions

### How do I generate LLVM IR from existing C++ source code?

Use the Clang compiler with the `-emit-llvm` flag to produce either human-readable assembly (`.ll`) or bitcode (`.bc`) files. For example, `clang -S -emit-llvm input.cpp` generates `input.ll` containing the textual IR representation that you can inspect or feed into custom optimization tools.

### What is the difference between the legacy PassManager and the new PassBuilder?

The legacy PassManager used implicit analysis invalidation and global pass registration, while the modern **PassBuilder** system—defined in [[`PassManager.h`](https://github.com/llvm/llvm-project/blob/main/PassManager.h)](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/PassManager.h)—uses explicit **analysis managers** and **preserved analyses** to track dependencies. This new design reduces memory overhead and enables finer-grained caching of analysis results across transformations.

### How do I add custom optimization passes to my pipeline?

Subclass `llvm::PassInfoMixin<YourPass>` and implement the `run()` method that accepts the unit of IR (e.g., `llvm::Function`) and an analysis manager. Then instantiate your pass and add it to the `FunctionPassManager` or `ModulePassManager` using `addPass()` before invoking the pipeline’s `run()` method.

### Can I optimize LLVM IR without using the C++ API?

Yes, you can use the `opt` command-line tool to run optimization passes on `.ll` or `.bc` files. However, for embedded compilers or JIT compilers, the C++ API provides deterministic control over which passes run and when, allowing you to tailor the optimization pipeline to specific hot paths or domain-specific requirements.