How Ghidra Handles Data Type Archives and .gdt Files: A Complete Technical Guide
Ghidra stores collections of data-type definitions in packed database archives with the .gdt extension, managed by the FileDataTypeManager class which provides version-aware loading, architecture metadata validation, and transaction-based persistence.
Ghidra, the NSA's open-source software reverse engineering framework, uses data type archives to organize and reuse type definitions across analysis projects. Understanding how these archives work is essential for extending Ghidra's type system or migrating legacy definitions. This article examines the internal mechanisms that power .gdt file handling, from the packed database format to the service layer that manages archive discovery and lifecycle.
What Are Data Type Archives and .gdt Files?
A .gdt file (Ghidra Data Types) is a packed database file that stores structured collections of data types. Unlike program-specific types embedded in a .gpr project, archives are standalone resources that can be shared across multiple reverse engineering sessions.
The archive format abstracts the underlying storage through a PackedDatabase accessed via PackedDBHandle. When you open a .gdt file, Ghidra does not load the entire archive into memory; instead, it establishes a managed connection to the packed database that supports lazy loading and transactional updates.
Core Architecture of .gdt File Handling
FileDataTypeManager and the Archive Implementation
The primary entry point for working with .gdt files is FileDataTypeManager, located in Ghidra/Framework/SoftwareModeling/src/main/java/ghidra/program/model/data/FileDataTypeManager.java. This class extends StandAloneDataTypeManager and implements the FileArchiveBasedDataTypeManager marker interface, providing concrete implementations for:
- Creating new empty archives
- Opening existing archives in read-only or update mode
- Persisting changes via save and save-as operations
- Managing the underlying
ResourceFilereference
StandAloneDataTypeManager (found in StandAloneDataTypeManager.java) supplies the base functionality for any non-project archive, including transaction handling and warning reporting. When you invoke FileDataTypeManager.createFileArchive() or openFileArchive(), you receive an instance that encapsulates both the database connection and the type management logic.
Architecture Metadata and Version Validation
When an archive is created with a specific program architecture, Ghidra stores three critical identifiers in a database string map:
LANGUAGE_IDCOMPILER_SPEC_IDLANGUAGE_VERSION
During initialization (initializeOtherAdapters in StandAloneDataTypeManager), the manager attempts to resolve these IDs against the current Ghidra installation. If the language or compiler spec cannot be found, the manager records an ArchiveWarning.LANGUAGE_NOT_FOUND or COMPILER_SPEC_NOT_FOUND status. If the language version has changed since the archive was created, the warning LANGUAGE_UPGRADE_REQUIRED is set.
You can query these warnings programmatically:
FileDataTypeManager dtm = FileDataTypeManager.openFileArchive(file, false);
ArchiveWarning warning = dtm.getWarning();
if (warning != ArchiveWarning.NONE) {
String details = dtm.getWarningMessage(true);
// Handle version mismatch
}
Discovering and Opening Data Type Archives
Built-in Archive Discovery
Ghidra ships with standard archives like generic_clib.gdt and windows_vs12_32.gdt. The DataTypeArchiveUtility class (in Ghidra/Features/Base/src/main/java/ghidra/app/plugin/core/datamgr/util/DataTypeArchiveUtility.java) scans the installation at startup using:
for (ResourceFile file :
Application.findFilesByExtensionInApplication(FileDataTypeManager.SUFFIX)) {
GHIDRA_ARCHIVES.put(file.getName(), file);
}
This utility also handles legacy name remapping. For example, if code requests the old generic_C_lib.gdt, the utility automatically redirects to generic_clib.gdt through its internal remapping table.
Opening Archives via the Service Layer
The DataTypeArchiveService (in DataTypeArchiveService.java) provides the high-level API for opening archives:
DataTypeArchiveService service = tool.getService(DataTypeArchiveService.class);
DataTypeManager dtm = service.openDataTypeArchive("generic_clib");
For direct file access, use openArchive(ResourceFile file, boolean acquireWriteLock):
acquireWriteLock = trueopens the archive inOpenMode.UPDATE(mutable)acquireWriteLock = falseopens inOpenMode.IMMUTABLE(read-only)
If you attempt to open an archive requiring a language upgrade in update mode, Ghidra throws a LanguageVersionException, forcing you to either upgrade the archive or open it read-only.
Creating, Saving, and Converting Archives
Creating New Archives
To create an empty archive without architecture binding:
File archiveFile = new File("/tmp/MyTypes.gdt");
FileDataTypeManager dtm = FileDataTypeManager.createFileArchive(archiveFile);
To create an archive tied to a specific processor architecture:
LanguageID langId = new LanguageID("x86:LE:32:default");
CompilerSpecID csId = new CompilerSpecID("gcc");
FileDataTypeManager dtm = FileDataTypeManager.createFileArchive(
archiveFile, langId, csId);
This validates the language and compiler spec IDs before recording them in the archive metadata.
Persistence Operations
FileDataTypeManager supports three save modes:
dtm.save()– Writes changes back to the original file in-placedtm.saveAs(newFile)– Creates a copy with a new universal ID and updates the internal file referencedtm.saveAs(newFile, newUniversalId)– Saves with a specific UUID for migration scenarios
All save operations clear the undo/redo stacks after writing to disk.
Archive Migration and Transformation
Ghidra provides the DataTypeArchiveTransformer (both GUI wizard and command-line tool) for complex migrations. This utility supports:
- Re-assigning universal IDs to resolve conflicts
- Updating program-architecture metadata to match new language versions
- Translating type IDs when merging archives
The transformer validates that source and destination files carry the .gdt suffix before processing.
Working with .gdt Files in Ghidra Scripts
Listing Categories from a Built-in Archive
import ghidra.app.services.DataTypeArchiveService;
import ghidra.program.model.data.*;
DataTypeArchiveService dts = state.getTool().getService(DataTypeArchiveService.class);
DataTypeManager dtm = dts.openDataTypeArchive("generic_clib");
println("Root categories in generic_clib.gdt:");
for (Category cat : dtm.getRootCategory().getCategories()) {
println(" " + cat.getName());
}
Creating and Populating a New Archive
import ghidra.program.model.data.*;
import java.io.File;
File out = new File("/tmp/MyArchive.gdt");
FileDataTypeManager dtm = FileDataTypeManager.createFileArchive(out);
Structure struct = dtm.createStructure("MyStruct");
struct.add(new DWordDataType(), "a", null);
struct.add(new ByteDataType(), "b", null);
dtm.endTransaction(dtm.startTransaction("Add struct"), true);
dtm.save();
println("Archive written to " + out.getAbsolutePath());
Summary
.gdtfiles are packed-database archives that store reusable Ghidra data-type definitions independently of specific programs.FileDataTypeManagerinFileDataTypeManager.javais the core class for creating, opening, and persisting these archives, extendingStandAloneDataTypeManagerfor transaction and warning handling.- Architecture metadata (language ID, compiler spec, version) is stored in a DB string map and validated on load, with structured
ArchiveWarningenums reporting mismatches. DataTypeArchiveServiceandDataTypeArchiveUtilityprovide unified discovery of built-in archives, legacy name resolution, and mode-aware opening (read-only vs. update).- Migration tools like
DataTypeArchiveTransformersupport version upgrades and UUID reassignment for long-term archive maintenance.
Frequently Asked Questions
What is the difference between a .gdt file and a program's local data types?
A .gdt file is a standalone archive that can be shared across multiple Ghidra projects, while local data types are embedded within a specific program database. Archives use FileDataTypeManager whereas program types are managed by ProgramDataTypeManager. You can import types from a .gdt archive into a program, but the archive remains an independent resource.
How do I resolve "Language Upgrade Required" warnings when opening an archive?
According to the StandAloneDataTypeManager implementation, this warning indicates the language version stored in the archive (LANGUAGE_VERSION key) differs from the current Ghidra installation. You must either open the archive read-only (OpenMode.IMMUTABLE) or use the DataTypeArchiveTransformer tool to migrate the archive to the current language version before opening it in update mode.
Can multiple Ghidra instances open the same .gdt file simultaneously?
No. When you open an archive with acquireWriteLock=true (update mode), FileDataTypeManager acquires exclusive access to the packed database. Other instances attempting to open the same file will receive a locking exception. For concurrent access, open the archive read-only in all instances, or copy the archive for independent modification.
Where does Ghidra store its built-in data type archives?
Built-in archives are distributed within the Ghidra installation directory and discovered at runtime by DataTypeArchiveUtility, which scans for files matching the .gdt suffix using Application.findFilesByExtensionInApplication(). These archives are cached in a static map (GHIDRA_ARCHIVES) keyed by filename for quick resolution during the openDataTypeArchive() call.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →