TensorRT vs ONNX vs TFLite for Model Deployment: A Technical Comparison
TensorRT optimizes ONNX models for NVIDIA GPUs using aggressive layer fusion and mixed precision, while TFLite converts TensorFlow graphs to lightweight flat-buffers for mobile CPUs and edge devices, with ONNX serving as the framework-agnostic interchange format.
The scutan90/DeepLearning-500-questions repository provides authoritative technical guidance on production deployment strategies, specifically in ch17_模型压缩、加速及移动端部署/第十七章_模型压缩、加速及移动端部署.md. Understanding the architectural differences between TensorRT, ONNX, and TensorFlow Lite (TFLite) is essential for selecting the correct inference engine for your deep learning deployment pipeline.
Understanding the Three Components
TensorRT (NVIDIA Inference Optimizer)
TensorRT is NVIDIA's proprietary SDK for high-performance deep learning inference on NVIDIA GPUs. According to the source analysis in 第十七章_模型压缩、加速及移动端部署.md, TensorRT accepts ONNX graphs exported from PyTorch, TensorFlow, Caffe, and MXNet, then performs aggressive graph optimizations including network pruning, layer fusion (e.g., Conv-BN-ReLU), and kernel auto-tuning for CUDA.
ONNX (Open Neural Network Exchange)
ONNX is not a runtime but an open standard format for representing machine learning models. It acts as the bridge between training frameworks and deployment runtimes like TensorRT. The repository notes that TensorRT specifically uses ONNX as its primary import format for cross-framework compatibility, allowing models trained in PyTorch or TensorFlow to be optimized for NVIDIA hardware without framework-specific conversion.
TensorFlow Lite (Mobile Inference Runtime)
TensorFlow Lite (TFLite) is Google's solution for deploying TensorFlow models on mobile, embedded, and IoT devices. As documented in 第十七章_模型压缩、加速及移动端部署.md, TFLite uses native .tflite flat-buffer files generated by the TensorFlow Lite Converter. It targets ARM CPUs, Android NNAPI, iOS Core ML, and edge TPUs, with a runtime footprint of approximately 500KB compared to TensorRT's ~10MB.
Model Format and Framework Compatibility
The primary distinction begins with input formats. TensorRT requires models in ONNX format (or NVIDIA's proprietary formats), enabling it to ingest models from PyTorch via torch.onnx.export(), TensorFlow via tf2onnx, and other frameworks. This is documented in the TensorRT 加速原理 section of the repository.
TFLite consumes .tflite flat-buffers created exclusively through the TensorFlow Lite Converter. While TensorFlow models convert natively, PyTorch models must first convert to ONNX, then to TensorFlow, then to TFLite—a more complex pipeline that underscores why TFLite is primarily a TensorFlow ecosystem tool.
Hardware Target and Optimization Strategy
TensorRT: NVIDIA GPU Specialization
According to 第十七章_模型压缩、加速及移动端部署.md, TensorRT is optimized exclusively for NVIDIA GPUs (desktop, data-center, and Jetson edge devices). It exploits FP16 and INT8 tensor cores through kernel-level optimizations including:
- Network pruning: Removing unused outputs and dead layers
- Layer fusion: Combining Conv-BN-ReLU into single kernels
- Kernel auto-tuning: Selecting optimal CUDA implementations for target GPU architectures
TFLite: Cross-Platform Edge Deployment
TFLite targets a broader hardware spectrum including mobile CPUs, GPUs via OpenGL/Vulkan, Android NNAPI, and iOS Core ML. Its optimization strategy focuses on:
- Post-training quantization: Converting FP32 weights to INT8 or FP16 with minimal accuracy loss
- Operator fusion: Combining Conv-ReLU operations during conversion
- FlatBuffer serialization: Zero-copy deserialization for fast model loading on memory-constrained devices
Precision Support and Quantization
Both frameworks support mixed precision, but with different implementations. TensorRT supports FP32, FP16, and INT8 via calibration-based quantization. The INT8 mode requires a calibration dataset to determine dynamic ranges for activations, as noted in the TensorRT 加速原理-INT8 section.
TFLite supports FP32, INT8 (full integer quantization), and FLOAT16 (experimental). It emphasizes post-training quantization that can be applied without retraining, making it accessible for mobile developers who need to reduce model size by 4x with minimal accuracy degradation.
Runtime Characteristics and API
TensorRT Runtime
The TensorRT runtime requires the CUDA libraries and NVIDIA drivers, resulting in a deployment footprint of approximately 10MB. It provides C++ and Python APIs, with additional integrations via torch-tensorrt and tf-trt for framework-specific workflows. The repository references TensorRT API 重建 for custom plugin development when standard operators are insufficient.
TFLite Runtime
TFLite offers a minimal runtime of approximately 500KB, ideal for Android APK size constraints. It provides APIs in Python (tf.lite.Interpreter), C++, Java/Kotlin (Android), and Swift/Objective-C (iOS). The Flex delegate allows fallback to TensorFlow runtime for unsupported operators, though this increases the binary size.
Practical Implementation Examples
Deploying with TensorRT via ONNX
The following example demonstrates building a TensorRT engine from an ONNX model, enabling FP16 precision for NVIDIA GPUs:
import tensorrt as trt
import pycuda.driver as cuda
import pycuda.autoinit
import numpy as np
# 1️⃣ Build engine from ONNX
TRT_LOGGER = trt.Logger(trt.Logger.WARNING)
builder = trt.Builder(TRT_LOGGER)
network = builder.create_network(1 << int(trt.NetworkDefinitionCreationFlag.EXPLICIT_BATCH))
parser = trt.OnnxParser(network, TRT_LOGGER)
with open("model.onnx", "rb") as f:
parser.parse(f.read())
# Enable FP16 & INT8 (requires calibration for INT8)
builder.max_batch_size = 1
builder.max_workspace_size = 1 << 30 # 1 GiB
builder.fp16_mode = True
engine = builder.build_cuda_engine(network)
# 2️⃣ Allocate buffers
context = engine.create_execution_context()
input_shape = (1, 3, 224, 224) # adjust to model
input_mem = cuda.mem_alloc(np.prod(input_shape) * np.float16().nbytes)
output_mem = cuda.mem_alloc(engine.get_binding_shape(1).volume() * np.float16().nbytes)
# 3️⃣ Inference
def infer(img_np):
img_np = img_np.astype(np.float16).ravel()
cuda.memcpy_htod(input_mem, img_np)
context.execute_v2([int(input_mem), int(output_mem)])
output = np.empty(engine.get_binding_shape(1), dtype=np.float16)
cuda.memcpy_dtoh(output, output_mem)
return output
Key source references: TensorRT optimisation principles and INT8 support are described in the TensorRT 加速原理 and TensorRT 如何优化重构模型 sections of ch17_模型压缩、加速及移动端部署/第十七章_模型压缩、加速及移动端部署.md in the scutan90/DeepLearning-500-questions repository.
Deploying with TFLite
The following example shows loading and running inference with a quantized TFLite model on mobile-friendly hardware:
import tensorflow as tf
import numpy as np
# 1️⃣ Load the .tflite model
interpreter = tf.lite.Interpreter(model_path="model.tflite")
interpreter.allocate_tensors()
# 2️⃣ Get input / output details
input_details = interpreter.get_input_details()
output_details = interpreter.get_output_details()
# 3️⃣ Prepare input (example: 224x224 RGB image)
input_shape = input_details[0]['shape']
input_data = np.random.rand(*input_shape).astype(np.float32)
# 4️⃣ Run inference
interpreter.set_tensor(input_details[0]['index'], input_data)
interpreter.invoke()
output_data = interpreter.get_tensor(output_details[0]['index'])
print(output_data)
Key source references: The repository's discussion of INT8 量化 for TensorFlow‑Lite shows that the same model can be quantised to a .tflite file for mobile deployment, as documented in 第十七章_模型压缩、加速及移动端部署.md.
Summary
- TensorRT is a GPU-centric inference engine that consumes ONNX models and performs aggressive kernel-level optimizations including layer fusion, precision calibration, and network pruning for NVIDIA hardware.
- TensorFlow Lite is a cross-platform runtime designed for mobile and embedded devices, utilizing .tflite flat-buffers with built-in quantization tools to achieve minimal binary size (~500KB).
- ONNX serves as the framework-agnostic bridge format, enabling models from PyTorch, TensorFlow, and other frameworks to be optimized by TensorRT, while TFLite primarily operates within the TensorFlow ecosystem.
- Choose TensorRT when deploying to NVIDIA GPUs requiring maximum throughput, and TFLite when targeting diverse mobile CPUs, Android NNAPI, or iOS Core ML with strict memory constraints.
Frequently Asked Questions
Can TensorRT run TFLite models directly?
No, TensorRT cannot directly execute TFLite flat-buffer files. TensorRT requires models in ONNX format or NVIDIA-specific formats. To deploy a TensorFlow model on TensorRT, you must first convert it to ONNX using tools like tf2onnx, then build a TensorRT engine from the ONNX graph as shown in the scutan90/DeepLearning-500-questions repository examples.
Is ONNX a runtime or just a model format?
ONNX is primarily an open standard format for representing machine learning models, not a runtime itself. It serves as a bridge between training frameworks and deployment engines like TensorRT. While the ONNX Runtime exists as a separate inference engine, the DeepLearning-500-questions documentation specifically highlights ONNX as the input format for TensorRT's optimization pipeline, enabling cross-framework compatibility for NVIDIA GPU deployment.
Which should I choose for NVIDIA Jetson edge devices?
For NVIDIA Jetson platforms (Nano, Xavier, Orin), TensorRT is the optimal choice. According to the deployment guidelines in ch17_模型压缩、加速及移动端部署/第十七章_模型压缩、加速及移动端部署.md, TensorRT leverages the CUDA cores and Tensor Cores on Jetson devices, supporting FP16 and INT8 precision for maximum performance. While TFLite can run on Jetson via CPU fallback, it cannot utilize the NVIDIA GPU acceleration that makes TensorRT essential for real-time inference on these edge devices.
How do I convert a PyTorch model to TFLite?
Converting a PyTorch model to TFLite requires an intermediate step through ONNX or TensorFlow. First, export your PyTorch model to ONNX using torch.onnx.export(). Then convert the ONNX model to TensorFlow using onnx-tf or similar tools. Finally, use the TensorFlow Lite Converter (tf.lite.TFLiteConverter) to generate the .tflite flat-buffer. This multi-step pipeline reflects the ecosystem differences highlighted in the DeepLearning-500-questions repository, where TFLite is tightly integrated with the TensorFlow ecosystem while ONNX provides the universal bridge for cross-framework deployment.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →