What Is LLVM IR and How to Optimize It Programmatically

LLVM IR is a language-agnostic, static single assignment (SSA) based intermediate representation that serves as the backbone of the LLVM compiler infrastructure, and you can optimize it programmatically using the PassBuilder API to construct custom transformation pipelines in C++.

LLVM IR (Intermediate Representation) is the core data format that powers the entire LLVM ecosystem. Stored in the llvm/llvm-project repository, this SSA-based representation enables front-ends like Clang to generate portable, type-safe code that can be aggressively optimized before target-specific code generation.

Core Architecture of LLVM IR

LLVM IR provides a universal abstraction layer between source languages and target architectures.

Modules, Functions, and Basic Blocks

The top-level container is the llvm::Module, which owns a collection of llvm::Function objects. Each function consists of a list of llvm::BasicBlock instances, forming a control-flow graph. According to the source in llvm/include/llvm/IR/Module.h, the module maintains a symbol table for global variables and function declarations.

SSA Form and Type System

Every instruction in LLVM IR produces a typed llvm::Value in strict SSA form, meaning each variable is defined exactly once. The llvm::IRBuilder class—defined in [llvm/include/llvm/IR/IRBuilder.h](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/IRBuilder.h)—automates the creation of these values while managing insertion points and type inference.

Generating LLVM IR Programmatically

You can produce LLVM IR either through compiler front-ends or direct API construction.

Emitting IR with Clang

The simplest method uses Clang to lower C/C++ into human-readable IR:

clang -S -emit-llvm -o output.ll input.c

This generates a .ll file containing the textual representation of the intermediate code.

Building IR with IRBuilder

For custom code generation, the llvm::IRBuilder template class provides a fluent API. As implemented in [llvm/include/llvm/IR/IRBuilder.h](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/IRBuilder.h), this utility handles instruction insertion, automatic naming, and debug location tracking.

The following example constructs a simple addition function:

#include "llvm/IR/IRBuilder.h"
#include "llvm/IR/Verifier.h"
#include "llvm/Support/raw_ostream.h"

int main() {
  llvm::LLVMContext Ctx;
  llvm::Module Mod("demo", Ctx);
  llvm::IRBuilder<> Builder(Ctx);

  // Define function type: int (int, int)
  auto *Int32Ty = Builder.getInt32Ty();
  auto *FuncTy = llvm::FunctionType::get(Int32Ty, {Int32Ty, Int32Ty}, false);
  auto *Func = llvm::Function::Create(FuncTy, 
                                      llvm::Function::ExternalLinkage,
                                      "add", Mod);

  // Name the arguments
  auto ArgIt = Func->arg_begin();
  ArgIt->setName("a");
  (++ArgIt)->setName("b");

  // Create entry block
  llvm::BasicBlock *Entry = llvm::BasicBlock::Create(Ctx, "entry", Func);
  Builder.SetInsertPoint(Entry);

  // Generate add instruction: %sum = add i32 %a, %b
  llvm::Value *Sum = Builder.CreateAdd(Func->getArg(0), 
                                       Func->getArg(1), "sum");
  Builder.CreateRet(Sum);

  // Verify and print
  llvm::verifyFunction(*Func);
  Mod.print(llvm::outs(), nullptr);
}

Programmatic Optimization Pipelines

Once you have generated IR, you can optimize it without invoking external tools by using LLVM’s modern pass management infrastructure.

The PassBuilder Infrastructure

The llvm::PassBuilder class—declared in [llvm/include/llvm/IR/PassManager.h](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/PassManager.h)—constructs optimization pipelines for both function-level and module-level transformations. Unlike the legacy pass manager, this system uses explicit analysis managers and preserved analysis tokens to minimize recomputation.

Constructing and Running Passes

The typical workflow involves instantiating a PassBuilder, registering analysis passes, and building pipelines that run transformations like InstCombinePass (instruction combining). The InstCombinePass header is located at [llvm/include/llvm/Transforms/Scalar/InstCombine.h](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/Transforms/Scalar/InstCombine.h) and performs peephole optimizations on algebraic identities.

Here is a complete example of running an optimization pipeline:

#include "llvm/IR/PassManager.h"
#include "llvm/IR/LegacyPassManagers.h"
#include "llvm/Transforms/Scalar/InstCombine.h"

int main() {
  // Assume 'Mod' is a valid llvm::Module from previous construction
  
  llvm::PassBuilder PB(Mod.getContext());

  // Build function-level pipeline
  llvm::FunctionPassManager<> FPM = PB.buildFunctionPassPipeline();
  FPM.addPass(llvm::InstCombinePass());

  // Run on each function
  for (auto &F : Mod) {
    FPM.run(F, llvm::FunctionAnalysisManager());
  }

  // Build and run module-level pipeline
  llvm::ModulePassManager MPM = PB.buildModulePassPipeline();
  MPM.run(Mod, llvm::ModuleAnalysisManager());
}

This approach allows you to compose specific optimizations—such as aggressive instruction combining, global variable optimization, and control-flow simplification—into a single programmatic execution.

Summary

Frequently Asked Questions

How do I generate LLVM IR from existing C++ source code?

Use the Clang compiler with the -emit-llvm flag to produce either human-readable assembly (.ll) or bitcode (.bc) files. For example, clang -S -emit-llvm input.cpp generates input.ll containing the textual IR representation that you can inspect or feed into custom optimization tools.

What is the difference between the legacy PassManager and the new PassBuilder?

The legacy PassManager used implicit analysis invalidation and global pass registration, while the modern PassBuilder system—defined in [PassManager.h](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/IR/PassManager.h)—uses explicit analysis managers and preserved analyses to track dependencies. This new design reduces memory overhead and enables finer-grained caching of analysis results across transformations.

How do I add custom optimization passes to my pipeline?

Subclass llvm::PassInfoMixin<YourPass> and implement the run() method that accepts the unit of IR (e.g., llvm::Function) and an analysis manager. Then instantiate your pass and add it to the FunctionPassManager or ModulePassManager using addPass() before invoking the pipeline’s run() method.

Can I optimize LLVM IR without using the C++ API?

Yes, you can use the opt command-line tool to run optimization passes on .ll or .bc files. However, for embedded compilers or JIT compilers, the C++ API provides deterministic control over which passes run and when, allowing you to tailor the optimization pipeline to specific hot paths or domain-specific requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →