# TensorRT vs ONNX vs TFLite for Model Deployment: A Technical Comparison

> Compare TensorRT, ONNX, and TFLite for model deployment. Understand NVIDIA GPU optimization, mobile CPU conversion, and framework-agnostic formats to choose the best solution.

- Repository: [scutan90/DeepLearning-500-questions](https://github.com/scutan90/DeepLearning-500-questions)
- Tags: deep-dive
- Published: 2026-03-06

---

**TensorRT optimizes ONNX models for NVIDIA GPUs using aggressive layer fusion and mixed precision, while TFLite converts TensorFlow graphs to lightweight flat-buffers for mobile CPUs and edge devices, with ONNX serving as the framework-agnostic interchange format.**

The `scutan90/DeepLearning-500-questions` repository provides authoritative technical guidance on production deployment strategies, specifically in `ch17_模型压缩、加速及移动端部署/第十七章_模型压缩、加速及移动端部署.md`. Understanding the architectural differences between TensorRT, ONNX, and TensorFlow Lite (TFLite) is essential for selecting the correct inference engine for your deep learning deployment pipeline.

## Understanding the Three Components

### TensorRT (NVIDIA Inference Optimizer)

**TensorRT** is NVIDIA's proprietary SDK for high-performance deep learning inference on NVIDIA GPUs. According to the source analysis in `第十七章_模型压缩、加速及移动端部署.md`, TensorRT accepts **ONNX** graphs exported from PyTorch, TensorFlow, Caffe, and MXNet, then performs aggressive graph optimizations including network pruning, layer fusion (e.g., Conv-BN-ReLU), and kernel auto-tuning for CUDA.

### ONNX (Open Neural Network Exchange)

**ONNX** is not a runtime but an open standard format for representing machine learning models. It acts as the bridge between training frameworks and deployment runtimes like TensorRT. The repository notes that TensorRT specifically uses ONNX as its primary import format for cross-framework compatibility, allowing models trained in PyTorch or TensorFlow to be optimized for NVIDIA hardware without framework-specific conversion.

### TensorFlow Lite (Mobile Inference Runtime)

**TensorFlow Lite (TFLite)** is Google's solution for deploying TensorFlow models on mobile, embedded, and IoT devices. As documented in `第十七章_模型压缩、加速及移动端部署.md`, TFLite uses native **.tflite** flat-buffer files generated by the TensorFlow Lite Converter. It targets ARM CPUs, Android NNAPI, iOS Core ML, and edge TPUs, with a runtime footprint of approximately 500KB compared to TensorRT's ~10MB.

## Model Format and Framework Compatibility

The primary distinction begins with input formats. **TensorRT** requires models in **ONNX** format (or NVIDIA's proprietary formats), enabling it to ingest models from PyTorch via `torch.onnx.export()`, TensorFlow via `tf2onnx`, and other frameworks. This is documented in the **TensorRT 加速原理** section of the repository.

**TFLite** consumes **.tflite** flat-buffers created exclusively through the TensorFlow Lite Converter. While TensorFlow models convert natively, PyTorch models must first convert to ONNX, then to TensorFlow, then to TFLite—a more complex pipeline that underscores why TFLite is primarily a TensorFlow ecosystem tool.

## Hardware Target and Optimization Strategy

### TensorRT: NVIDIA GPU Specialization

According to `第十七章_模型压缩、加速及移动端部署.md`, TensorRT is optimized exclusively for **NVIDIA GPUs** (desktop, data-center, and Jetson edge devices). It exploits **FP16** and **INT8** tensor cores through kernel-level optimizations including:

- **Network pruning**: Removing unused outputs and dead layers
- **Layer fusion**: Combining Conv-BN-ReLU into single kernels
- **Kernel auto-tuning**: Selecting optimal CUDA implementations for target GPU architectures

### TFLite: Cross-Platform Edge Deployment

**TFLite** targets a broader hardware spectrum including **mobile CPUs**, **GPUs via OpenGL/Vulkan**, **Android NNAPI**, and **iOS Core ML**. Its optimization strategy focuses on:

- **Post-training quantization**: Converting FP32 weights to INT8 or FP16 with minimal accuracy loss
- **Operator fusion**: Combining Conv-ReLU operations during conversion
- **FlatBuffer serialization**: Zero-copy deserialization for fast model loading on memory-constrained devices

## Precision Support and Quantization

Both frameworks support mixed precision, but with different implementations. **TensorRT** supports **FP32**, **FP16**, and **INT8** via calibration-based quantization. The INT8 mode requires a calibration dataset to determine dynamic ranges for activations, as noted in the **TensorRT 加速原理-INT8** section.

**TFLite** supports **FP32**, **INT8** (full integer quantization), and **FLOAT16** (experimental). It emphasizes post-training quantization that can be applied without retraining, making it accessible for mobile developers who need to reduce model size by 4x with minimal accuracy degradation.

## Runtime Characteristics and API

### TensorRT Runtime

The TensorRT runtime requires the **CUDA libraries** and NVIDIA drivers, resulting in a deployment footprint of approximately **10MB**. It provides C++ and Python APIs, with additional integrations via **torch-tensorrt** and **tf-trt** for framework-specific workflows. The repository references **TensorRT API 重建** for custom plugin development when standard operators are insufficient.

### TFLite Runtime

TFLite offers a minimal runtime of approximately **500KB**, ideal for Android APK size constraints. It provides APIs in Python (`tf.lite.Interpreter`), C++, Java/Kotlin (Android), and Swift/Objective-C (iOS). The **Flex delegate** allows fallback to TensorFlow runtime for unsupported operators, though this increases the binary size.

## Practical Implementation Examples

### Deploying with TensorRT via ONNX

The following example demonstrates building a TensorRT engine from an ONNX model, enabling FP16 precision for NVIDIA GPUs:

```python
import tensorrt as trt
import pycuda.driver as cuda
import pycuda.autoinit
import numpy as np

# 1️⃣ Build engine from ONNX

TRT_LOGGER = trt.Logger(trt.Logger.WARNING)
builder = trt.Builder(TRT_LOGGER)
network = builder.create_network(1 << int(trt.NetworkDefinitionCreationFlag.EXPLICIT_BATCH))
parser = trt.OnnxParser(network, TRT_LOGGER)

with open("model.onnx", "rb") as f:
    parser.parse(f.read())

# Enable FP16 & INT8 (requires calibration for INT8)

builder.max_batch_size = 1
builder.max_workspace_size = 1 << 30   # 1 GiB

builder.fp16_mode = True
engine = builder.build_cuda_engine(network)

# 2️⃣ Allocate buffers

context = engine.create_execution_context()
input_shape = (1, 3, 224, 224)               # adjust to model

input_mem = cuda.mem_alloc(np.prod(input_shape) * np.float16().nbytes)
output_mem = cuda.mem_alloc(engine.get_binding_shape(1).volume() * np.float16().nbytes)

# 3️⃣ Inference

def infer(img_np):
    img_np = img_np.astype(np.float16).ravel()
    cuda.memcpy_htod(input_mem, img_np)
    context.execute_v2([int(input_mem), int(output_mem)])
    output = np.empty(engine.get_binding_shape(1), dtype=np.float16)
    cuda.memcpy_dtoh(output, output_mem)
    return output

```

*Key source references*: TensorRT optimisation principles and INT8 support are described in the **TensorRT 加速原理** and **TensorRT 如何优化重构模型** sections of `ch17_模型压缩、加速及移动端部署/第十七章_模型压缩、加速及移动端部署.md` in the `scutan90/DeepLearning-500-questions` repository.

### Deploying with TFLite

The following example shows loading and running inference with a quantized TFLite model on mobile-friendly hardware:

```python
import tensorflow as tf
import numpy as np

# 1️⃣ Load the .tflite model

interpreter = tf.lite.Interpreter(model_path="model.tflite")
interpreter.allocate_tensors()

# 2️⃣ Get input / output details

input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()

# 3️⃣ Prepare input (example: 224x224 RGB image)

input_shape = input_details[0]['shape']
input_data = np.random.rand(*input_shape).astype(np.float32)

# 4️⃣ Run inference

interpreter.set_tensor(input_details[0]['index'], input_data)
interpreter.invoke()
output_data = interpreter.get_tensor(output_details[0]['index'])
print(output_data)

```

*Key source references*: The repository's discussion of **INT8 量化** for TensorFlow‑Lite shows that the same model can be quantised to a `.tflite` file for mobile deployment, as documented in `第十七章_模型压缩、加速及移动端部署.md`.

## Summary

- **TensorRT** is a GPU-centric inference engine that consumes **ONNX** models and performs aggressive kernel-level optimizations including layer fusion, precision calibration, and network pruning for NVIDIA hardware.
- **TensorFlow Lite** is a cross-platform runtime designed for mobile and embedded devices, utilizing **.tflite** flat-buffers with built-in quantization tools to achieve minimal binary size (~500KB).
- **ONNX** serves as the framework-agnostic bridge format, enabling models from PyTorch, TensorFlow, and other frameworks to be optimized by TensorRT, while TFLite primarily operates within the TensorFlow ecosystem.
- Choose **TensorRT** when deploying to NVIDIA GPUs requiring maximum throughput, and **TFLite** when targeting diverse mobile CPUs, Android NNAPI, or iOS Core ML with strict memory constraints.

## Frequently Asked Questions

### Can TensorRT run TFLite models directly?

No, TensorRT cannot directly execute TFLite flat-buffer files. TensorRT requires models in **ONNX** format or NVIDIA-specific formats. To deploy a TensorFlow model on TensorRT, you must first convert it to ONNX using tools like `tf2onnx`, then build a TensorRT engine from the ONNX graph as shown in the `scutan90/DeepLearning-500-questions` repository examples.

### Is ONNX a runtime or just a model format?

**ONNX** is primarily an open standard format for representing machine learning models, not a runtime itself. It serves as a bridge between training frameworks and deployment engines like TensorRT. While the ONNX Runtime exists as a separate inference engine, the `DeepLearning-500-questions` documentation specifically highlights ONNX as the input format for TensorRT's optimization pipeline, enabling cross-framework compatibility for NVIDIA GPU deployment.

### Which should I choose for NVIDIA Jetson edge devices?

For **NVIDIA Jetson** platforms (Nano, Xavier, Orin), **TensorRT** is the optimal choice. According to the deployment guidelines in `ch17_模型压缩、加速及移动端部署/第十七章_模型压缩、加速及移动端部署.md`, TensorRT leverages the CUDA cores and Tensor Cores on Jetson devices, supporting **FP16** and **INT8** precision for maximum performance. While TFLite can run on Jetson via CPU fallback, it cannot utilize the NVIDIA GPU acceleration that makes TensorRT essential for real-time inference on these edge devices.

### How do I convert a PyTorch model to TFLite?

Converting a PyTorch model to TFLite requires an intermediate step through ONNX or TensorFlow. First, export your PyTorch model to **ONNX** using `torch.onnx.export()`. Then convert the ONNX model to TensorFlow using `onnx-tf` or similar tools. Finally, use the **TensorFlow Lite Converter** (`tf.lite.TFLiteConverter`) to generate the `.tflite` flat-buffer. This multi-step pipeline reflects the ecosystem differences highlighted in the `DeepLearning-500-questions` repository, where TFLite is tightly integrated with the TensorFlow ecosystem while ONNX provides the universal bridge for cross-framework deployment.