How to Migrate from llama3.java to GPULlama3.java for GPU Acceleration
Migrating from llama3.java to GPULlama3.java requires swapping the Maven dependency to io.github.beehive-lab:gpu-llama3, installing the TornadoVM SDK, and enabling GPU execution via the --use-tornadovm CLI flag or the .onGPU(Boolean.TRUE) builder method.
The GPULlama3.java repository (beehive-lab/gpullama3.java) provides a drop-in replacement for the original CPU-only llama3.java that accelerates transformer inference using TornadoVM kernels. While the tokenizer, GGUF loader, and sampling logic remain identical in the org.beehive.gpullama3 package, the inference engine delegates to GPU compute kernels when explicitly enabled.
Why Migrate to GPULlama3.java?
GPULlama3.java preserves the exact Java-first API of the original project while adding TornadoVM-based GPU execution. This means you can migrate existing inference code without rewriting tokenization or sampling logic.
Key benefits include:
- Hardware acceleration: Transformer layers execute on NVIDIA, Intel, and AMD GPUs via OpenCL and PTX backends
- Identical public API: The
GPULlama3ChatModelclass maintains the same builder pattern as the originalLlamaChatModel - Seamless fallback: When
useTornadovmis false, the system automatically uses the standard CPU implementation inInferenceEngine
Key Architectural Changes
Understanding the internal differences helps diagnose migration issues. The core logic remains in the org.beehive.gpullama3 package, but specific components swap implementations based on the useTornadovm flag:
| Component | CPU Implementation (llama3.java) | GPU Implementation (GPULlama3.java) | Source File |
|---|---|---|---|
| Inference Engine | Standard Java arrays | TornadoVM kernels (TransformerComputeKernels) |
TransformerComputeKernels.java |
| Weight Storage | StandardWeights (float arrays) |
TornadoWeights (GPU memory buffers) |
TornadoWeights.java |
| Model Loading | AbstractModelLoader |
LlamaModelLoader with useTornadovm flag |
ModelType.java |
| Configuration | No GPU support | --use-tornadovm flag or onGPU() builder method |
Options.java |
When useTornadovm is set to true in ModelType.java, the loader instantiates LlamaTornadoWeights, which allocate device memory and compile TornadoVM kernels for the transformer layers.
Step-by-Step Migration Guide
Update Your Dependencies
Replace the original llama3.java artifact with the GPULlama3.java dependency in your pom.xml:
<dependency>
<groupId>io.github.beehive-lab</groupId>
<artifactId>gpu-llama3</artifactId>
<version>0.4.0</version>
</dependency>
Choose the variant matching your JDK version (JDK 21 or JDK 25).
Install TornadoVM
GPULlama3.java requires the TornadoVM native runtime. Install it using the SDK manager or manual download:
sdk install tornadovm
tornado --devices
The tornado --devices command verifies that your GPU is detected and compatible.
Enable GPU via Command Line
If you currently run inference using the CLI, add the --use-tornadovm flag:
Before (CPU):
java -jar llama3.jar --model mymodel.gguf --prompt "Hello"
After (GPU):
java -jar gpulambda3.jar --model mymodel.gguf --prompt "Hello" --use-tornadovm true
As defined in Options.java, the parser checks for --use-tornadovm and falls back to the system property use.tornadovm if the flag is absent.
Enable GPU via Builder API
For programmatic access, update your imports and use the builder method:
import org.beehive.gpullama3.GPULlama3ChatModel;
import java.nio.file.Paths;
public class MigrationExample {
public static void main(String[] args) {
GPULlama3ChatModel model = GPULlama3ChatModel.builder()
.modelPath(Paths.get("models/beehive-llama-3.2-1b-instruct-fp16.gguf"))
.temperature(0.9f)
.topP(0.9f)
.maxTokens(2048)
.onGPU(Boolean.TRUE) // Enable GPU acceleration
.build();
String response = model.generate("Explain GPU acceleration");
System.out.println(response);
}
}
The onGPU(Boolean.TRUE) setting propagates through to LlamaModelLoader, triggering the instantiation of TornadoWeights and the compilation of kernels in TransformerComputeKernels.java.
Verifying GPU Acceleration
Confirm that TornadoVM is active by using the --show-command flag:
llama-tornado --gpu --model beehive-llama-3.2-1b-instruct-fp16.gguf \
--prompt "Test" --show-command
This prints the full JVM launch command, which should include -Dtornado.device=<backend> and memory allocation flags, confirming that GPU kernels will execute the transformer layers rather than CPU loops.
Troubleshooting Common Migration Issues
| Symptom | Root Cause | Solution |
|---|---|---|
| "No GPU devices found" | TornadoVM not installed or JAVA_HOME points to unsupported JDK |
Install TornadoVM SDK matching your JDK version and verify with tornado --devices |
UnsupportedOperationException from ModelType.UNKNOWN |
Invalid GGUF path or corrupted model file | Ensure the --model path points to a valid GGUF file; the loader auto-detects the model type |
| GPU out-of-memory | Default 7GB allocation insufficient for model | Increase allocation with --gpu-memory 12GB |
| Silent CPU fallback | Hardware lacks PTX/OpenCL support (e.g., Apple Silicon without proper drivers) | Verify GPU compatibility in the TornadoVM README; check supported backends (NVIDIA, Intel, AMD) |
Summary
- GPULlama3.java is a drop-in replacement for llama3.java that accelerates inference via TornadoVM while maintaining the same public API
- Migration requires updating the Maven dependency to
io.github.beehive-lab:gpu-llama3and installing the TornadoVM SDK - Enable GPU execution using either the
--use-tornadovm trueCLI flag (parsed inOptions.java) or the.onGPU(Boolean.TRUE)builder method - GPU weights are managed by
TornadoWeights.java, while compute kernels reside inTransformerComputeKernels.java - All tokenizer, sampler, and GGUF loading code remains unchanged during migration
Frequently Asked Questions
Do I need to rewrite my existing llama3.java code to use GPULlama3.java?
No. GPULlama3.java maintains API compatibility with the original project. You only need to change the import from org.beehive.llama3.LlamaChatModel to org.beehive.gpullama3.GPULlama3ChatModel and add the .onGPU(Boolean.TRUE) configuration. The tokenizer, sampler, and prompt handling logic remain identical.
What happens if TornadoVM is not installed but I set onGPU(true)?
The application will fail to initialize with an error indicating that TornadoVM cannot be detected. Unlike silent failures, the ModelType factory and LlamaModelLoader explicitly validate the useTornadovm flag and require the TornadoVM runtime to be present in JAVA_HOME when GPU mode is requested.
Can I switch between CPU and GPU execution without changing code?
Yes. When using the CLI, omitting --use-tornadovm or setting it to false forces the InferenceEngine to use standard Java arrays and StandardWeights instead of TornadoWeights. When using the builder API, simply pass Boolean.FALSE or omit the onGPU() call to default to CPU execution.
Which GPUs are supported by GPULlama3.java?
GPULlama3.java supports NVIDIA GPUs via PTX, and Intel/AMD GPUs via OpenCL. The specific compatibility depends on your TornadoVM installation. Run tornado --devices to list available backends. Apple Silicon support varies based on OpenCL driver availability.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →