# How LLVM Bitcode Format Works for Serialization: A Deep Dive into IR Module Encoding

> Discover how LLVM bitcode serializes IR modules with a hierarchical block and record structure. Learn about efficient storage, transmission, and lazy loading.

- Repository: [LLVM/llvm-project](https://github.com/llvm/llvm-project)
- Tags: deep-dive
- Published: 2026-09-11

---

**LLVM bitcode is a compact binary format that serializes LLVM IR modules using a hierarchical block-and-record structure built on the Bitstream library, enabling efficient storage, transmission, and lazy loading.**

The llvm/llvm-project repository implements this format to persist intermediate representation (IR) across compiler stages. Understanding how LLVM bitcode serialization functions is essential for tool developers working with link-time optimization, distributed builds, or custom LLVM-based compilers.

## Architecture of LLVM Bitcode Serialization

LLVM bitcode operates as a layered encoding system. The format rests on three foundational layers: the bitstream transport layer, the semantic encoder/decoder layer, and the opcode definition layer.

### Bitstream Foundation

At the lowest level, the **Bitstream** library handles raw binary encoding. The `BitstreamWriter` class in [`llvm/lib/Bitstream/BitstreamWriter.h`](https://github.com/llvm/llvm-project/blob/main/llvm/lib/Bitstream/BitstreamWriter.h) writes sequences of 32-bit words, while `BitstreamReader` decodes them. This layer manages bit-level packing, ensuring integers and strings occupy minimal space without requiring byte alignment.

### Writer and Reader Components

The semantic layer resides in [`llvm/lib/Bitcode/Writer/BitcodeWriter.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm/lib/Bitcode/Writer/BitcodeWriter.cpp) and [`llvm/lib/Bitcode/Reader/BitcodeReader.cpp`](https://github.com/llvm/llvm-project/blob/main/llvm/lib/Bitcode/Reader/BitcodeReader.cpp). 

- **BitcodeWriter**: Traverses an `llvm::Module`, assigns numeric IDs via `ValueEnumerator`, and emits hierarchical blocks. The entry point `write()` method orchestrates the serialization pipeline around line 22-24 of [`BitcodeWriter.cpp`](https://github.com/llvm/llvm-project/blob/main/BitcodeWriter.cpp).
- **BitcodeReader**: Parses the bitstream using `BitstreamCursor`, reconstructs type tables, and lazily materializes functions. The `ParseBitcodeFile` method validates headers and iteratively reconstructs the module graph.

### Opcode Definitions

All block and record identifiers live in [`llvm/include/llvm/Bitcode/LLVMBitCodes.h`](https://github.com/llvm/llvm-project/blob/main/llvm/include/llvm/Bitcode/LLVMBitCodes.h). This header defines constants like `MODULE_CODE_VERSION`, `TYPE_CODE_FUNCTION`, and `FUNC_CODE_DECLAREBLOCKS`, ensuring the writer and reader maintain contract compatibility.

## Block Hierarchy and Record Structure

LLVM bitcode organizes data into nested blocks. Every file begins with a module block (`bitc::MODULE_BLOCK_ID`), which contains specialized sub-blocks for different IR entities.

### Module Block and Sub-blocks

The writer creates distinct sub-blocks using `Stream.EnterSubblock(BlockID, Revision)` and closes them with `Stream.ExitBlock()`:

- **TYPE_BLOCK_ID**: Stores all LLVM types as `TYPE_CODE_*` records. Each function signature, struct, and primitive type receives a sequential identifier.
- **PARAMATTR_BLOCK_ID**: Contains parameter attribute groups accessed via `PARAMATTR_GRP_CODE_ENTRY` codes. The writer populates this using `writeAttributeGroupTable` and `writeAttributeTable`.
- **CONST_BLOCK_ID**: Serializes constants (integers, floats, aggregates) using codes like `CONST_CODE_INTEGER`.
- **FUNCTION_BLOCK_ID**: Encodes individual functions with their basic blocks and instructions. The `FUNC_CODE_DECLAREBLOCKS` record marks the start of a function body.
- **METADATA_BLOCK_ID**: Holds debug information and named metadata via `METADATA_CODE_VERSION` and related records.

### Abbreviations and Compact Encoding

Within each block, records are emitted via `Stream.EmitRecord(Code, Operands, Abbrev)`. The optional `Abbrev` parameter enables custom compact encodings. The writer pre-defines abbreviations (see lines 71-78 of [`BitcodeWriter.cpp`](https://github.com/llvm/llvm-project/blob/main/BitcodeWriter.cpp)) to reduce bitcode size for common patterns like small integers or frequent instruction types.

## Value Enumeration and Type Mapping

Before emitting records, the writer constructs a **value enumerator** (`ValueEnumerator VE`) to eliminate redundancy.

### ValueEnumerator Implementation

The `ValueEnumerator` assigns sequential IDs to every function, global variable, constant, and type. Key methods include:
- `VE.getValueID()`: Retrieves the numeric identifier for values
- `VE.getTypeID()`: Retrieves the numeric identifier for types

These IDs replace pointer references throughout the bitstream, allowing the reader to reconstruct the original graph without duplication.

### Attribute Group Serialization

The enumerator also collects attribute groups and lists via `VE.getAttributeGroups()` and `VE.getAttributeLists()`. The writer serializes these into `PARAMATTR` blocks before encoding the values that reference them, ensuring the reader can reconstruct function signatures completely.

## Metadata and Debug Information Encoding

Metadata receives special handling in `METADATA_BLOCK_ID`. The writer converts `DILocation`, `DICompileUnit`, and other debug nodes into integer and string sequences using specialized methods like `writeDILocation` and `writeDICompileUnit`. 

The reader reverses this process, reconstructing `llvm::Metadata` objects and attaching them to the appropriate IR values. This preserves source-level debugging information through the serialization round-trip.

## Versioning and Forward Compatibility

Every bitcode file begins with a version record (`MODULE_CODE_VERSION`) that stores the module version number. The reader validates this through `hasInvalidBitcodeHeader` and `readIdentificationBlock` to ensure producer-consumer compatibility.

Unknown record codes are safely skipped during parsing, allowing newer LLVM versions to extend the format without breaking older readers. This forward-compatibility mechanism ensures longevity for bitcode files in production environments.

## Lazy Loading and Thin LTO Support

LLVM bitcode supports incremental deserialization for large modules. The `BitcodeReader` can materialize functions on-demand rather than loading the entire module into memory.

### ThinLTO Index Files

For Thin Link-Time Optimization (ThinLTO), the `IndexBitcodeWriter` generates specialized bitcode containing only summary information—specifically GUID-to-value-ID mappings—rather than full function definitions. This index format enables distributed compilation workflows where only changed modules require full recompilation.

### Use-List Order Preservation

When the `preserve-bc-uselistorder` flag is enabled, the writer emits a **use-list order** block. This ensures deterministic reconstruction of instruction ordering, which is critical for reproducible builds and certain optimization passes.

## Practical Implementation

The following examples demonstrate end-to-end serialization using the LLVM C++ API.

### Serializing a Module to Bitcode

```cpp
#include "llvm/IR/Module.h"
#include "llvm/IR/LLVMContext.h"
#include "llvm/Bitcode/BitcodeWriter.h"
#include "llvm/Support/raw_ostream.h"

int main() {
  llvm::LLVMContext Ctx;
  // Build a trivial module: `define i32 @foo() { ret i32 42 }`
  std::unique_ptr<llvm::Module> M = std::make_unique<llvm::Module>("demo", Ctx);
  llvm::FunctionType *FT = llvm::FunctionType::get(llvm::Type::getInt32Ty(Ctx), {}, false);
  llvm::Function *F = llvm::Function::Create(FT, llvm::GlobalValue::ExternalLinkage, "foo", *M);
  llvm::BasicBlock *BB = llvm::BasicBlock::Create(Ctx, "entry", F);
  llvm::IRBuilder<> Builder(BB);
  Builder.CreateRet(llvm::ConstantInt::get(llvm::Type::getInt32Ty(Ctx), 42));

  // Write the module to a file
  std::error_code EC;
  llvm::raw_fd_ostream OS("demo.bc", EC, llvm::sys::fs::OF_None);
  llvm::WriteBitcodeToFile(*M, OS);
  return EC ? 1 : 0;
}

```

### Deserializing Bitcode to a Module

```cpp
#include "llvm/IR/LLVMContext.h"
#include "llvm/Bitcode/BitcodeReader.h"
#include "llvm/Support/MemoryBuffer.h"
#include "llvm/Support/raw_ostream.h"

int main() {
  llvm::LLVMContext Ctx;
  // Load the bitcode file into a MemoryBuffer
  auto BufferOrError = llvm::MemoryBuffer::getFile("demo.bc");
  if (!BufferOrError) return 1;
  std::unique_ptr<llvm::MemoryBuffer> Buffer = std::move(*BufferOrError);

  // Parse the bitcode into a Module
  auto ModuleOrError = llvm::parseBitcodeFile(Buffer->getMemBufferRef(), Ctx);
  if (!ModuleOrError) return 1;
  std::unique_ptr<llvm::Module> M = std::move(ModuleOrError.get());

  // Print the textual IR to stdout
  M->print(llvm::outs(), nullptr);
  return 0;
}

```

## Summary

- LLVM bitcode uses a **hierarchical block structure** (MODULE, TYPE, FUNCTION, etc.) defined in [`LLVMBitCodes.h`](https://github.com/llvm/llvm-project/blob/main/LLVMBitCodes.h) to organize IR entities.
- The **Bitstream** library ([`BitstreamWriter.h`](https://github.com/llvm/llvm-project/blob/main/BitstreamWriter.h)) provides low-level bit packing, while [`BitcodeWriter.cpp`](https://github.com/llvm/llvm-project/blob/main/BitcodeWriter.cpp) and [`BitcodeReader.cpp`](https://github.com/llvm/llvm-project/blob/main/BitcodeReader.cpp) handle semantic encoding.
- **ValueEnumerator** eliminates redundancy by assigning numeric IDs to all values and types before serialization.
- **Abbreviations** enable compact encoding for frequently occurring patterns, reducing file size.
- **Version records** and safe skipping of unknown codes ensure forward compatibility across LLVM versions.
- **ThinLTO** utilizes specialized index bitcode containing only GUID mappings for distributed optimization.

## Frequently Asked Questions

### What is the difference between LLVM bitcode and LLVM IR?

LLVM IR is the textual intermediate representation that humans read and write, while LLVM bitcode is the compact binary serialization of that same IR. The bitcode format uses numeric codes and compressed structures to represent the identical semantic information in a machine-optimized format suitable for storage and rapid loading.

### How does LLVM bitcode achieve compact file sizes?

The format employs multiple compression strategies: the **Bitstream** layer packs data at the bit level rather than byte alignment, **ValueEnumerator** eliminates duplicate value references through numeric IDs, and **abbreviations** allow custom compact encodings for common integer ranges and record patterns. These mechanisms work together in [`BitcodeWriter.cpp`](https://github.com/llvm/llvm-project/blob/main/BitcodeWriter.cpp) to minimize output size.

### Can older LLVM versions read bitcode from newer versions?

Generally yes, due to the format's **forward compatibility** design. Each block begins with a version record that the reader validates via `hasInvalidBitcodeHeader`. When the reader in [`BitcodeReader.cpp`](https://github.com/llvm/llvm-project/blob/main/BitcodeReader.cpp) encounters unknown record codes, it skips them safely rather than failing. However, semantic changes to IR constructs may still cause incompatibility for specific features.

### How does ThinLTO use LLVM bitcode differently?

Standard bitcode contains complete function definitions, but ThinLTO generates **index bitcode** files containing only summary information. The `IndexBitcodeWriter` class emits GUID-to-value-ID mappings and global summaries without full function bodies. This allows the linker to perform whole-program analysis while deferring full parsing and optimization until the backend stage.