How ncnn Manages Model Weights: load_model() vs External Memory Referencing
ncnn manages model weights through an abstraction layer called ModelBin, allowing the load_model() method to ingest weights from disk files, memory buffers, or custom data streams without requiring layers to distinguish between the sources.
In the Tencent/ncnn inference framework, model parameters—including weights, biases, and quantization tables—are decoupled from the network architecture to support flexible deployment scenarios. The framework provides multiple overloads of ncnn::Net::load_model in src/net.cpp to handle different data sources, ranging from traditional file I/O to zero-copy external memory referencing. Understanding how ncnn manages model weights through these pathways is essential for optimizing memory usage and loading latency in production environments.
The ModelBin Abstraction Layer
All layers in ncnn retrieve their weights through the ModelBin interface, defined in include/ncnn/modelbin.h. This abstraction ensures that layer implementations remain agnostic to the underlying storage mechanism.
When a network loads weights, each layer's load_model method receives a const ncnn::ModelBin& reference. The layer then calls mb.load(index) to retrieve tensors:
int Convolution::load_model(const ncnn::ModelBin& mb)
{
weight_data = mb.load(w);
bias_data = mb.load(b);
// ...
}
The ModelBin implementation handles the actual byte reading. ModelBinFromMatArray wraps a raw pointer to user-supplied memory (such as a std::vector<float>), while the default implementation uses a DataReader for file-based access. This design allows the same layer code to function correctly whether the weights reside on disk or in RAM.
Net::load_model() Overloads Explained
The ncnn::Net class provides four primary overloads of load_model() in src/net.cpp, all converging on the same internal loop between lines 1600–1700 that iterates over layers and invokes layer->load_model(mb).
File-Based Loading (Standard I/O)
For deployment scenarios where the model binary (.bin) resides on the filesystem, ncnn offers two file-based overloads:
load_model(const char* modelpath)(src/net.cpp#L1838): Opens the specified path as aDataReaderFile.load_model(FILE* fp)(src/net.cpp#L1832): Wraps an existingFILE*stream in aDataReaderFile.
These methods are suitable for desktop applications and standard mobile deployments where the model file is packaged as an app asset.
Memory-Based Loading (External Referencing)
When the weight data is already present in memory—perhaps decompressed from an archive, received over a network, or stored in a memory-mapped file—the load_model(const unsigned char* mem) overload (src/net.cpp#L1923) provides a zero-copy pathway. This method constructs a ModelBinFromMatArray instance that points directly to the provided buffer without allocating internal copies of the weight tensors.
Custom Data Sources
For encrypted or compressed models, developers can subclass DataReader and pass it to load_model(const DataReader& dr) (src/net.cpp#L1599). This central routine builds the ModelBin from the reader and executes the layer loading loop, enabling on-the-fly decryption or streaming from non-standard sources.
How External Memory Referencing Works
External memory referencing leverages ModelBinFromMatArray, implemented in src/mat.cpp and declared in include/ncnn/modelbin.h. When you invoke the pointer overload:
const unsigned char* model_data = /* pre-loaded blob */;
ncnn::Net net;
net.load_model(model_data);
The framework treats model_data as a read-only array of floats and constructs weight matrices that reference this memory directly. This approach eliminates the memory overhead of duplicating large weight files and avoids file system latency during initialization.
The unit tests in tests/test_squeezenet.cpp demonstrate this pattern in lines 179–201, where the test suite exercises load_model((const unsigned char*)model_data) to validate inference without touching the disk.
Performance Considerations and Use Cases
Choosing the appropriate load_model overload depends on your deployment constraints and performance requirements:
- Standard file deployment: Use
load_model(const char*)for simplicity when the.binfile is stored on disk or in standard Android/iOS assets. - Embedded systems with custom loaders: Use
load_model(const unsigned char*)when the model is embedded as a byte array in the executable or loaded into RAM by a custom resource manager. - Security-sensitive applications: Implement a custom
DataReadersubclass forload_model(const DataReader&)when models are encrypted or compressed, allowing decryption to occur stream-wise during layer initialization. - Low-latency inference: The external memory pointer overload minimizes initialization time and peak memory usage by avoiding file I/O and buffer copies, making it ideal for real-time applications on resource-constrained devices.
Practical Implementation Examples
Loading Weights from Disk
The most common pattern loads network parameters first, then the binary weights:
#include "net.h"
ncnn::Net net;
net.load_param("mobilenet_v2.param");
net.load_model("mobilenet_v2.bin");
Source: Net::load_model(const char*) (src/net.cpp#L1838).
Loading from External Memory
For zero-copy loading when the binary data is already in a std::vector or mmap buffer:
#include "net.h"
#include <vector>
std::vector<unsigned char> buffer = read_from_custom_source();
ncnn::Net net;
net.load_param("shufflenet_v2.param");
net.load_model(buffer.data()); // No copy performed
Source: Net::load_model(const unsigned char*) (src/net.cpp#L1923).
Custom DataReader for Encrypted Models
To load weights from an encrypted blob without writing a temporary file:
class EncryptedReader : public ncnn::DataReader {
public:
EncryptedReader(const unsigned char* key, const unsigned char* data, size_t len)
: key_(key), data_(data), pos_(0), len_(len) {}
virtual int read(void* buf, size_t size) {
decrypt_block(data_ + pos_, buf, size, key_);
pos_ += size;
return 0;
}
virtual void close() {}
private:
const unsigned char *key_, *data_;
size_t pos_, len_;
};
// Usage
EncryptedReader reader(key, encrypted_blob, blob_size);
net.load_model(reader);
Source: Net::load_model(const DataReader&) (src/net.cpp#L1599).
Summary
- ncnn uses the
ModelBinabstraction to decouple layer weight consumption from storage implementations, located ininclude/ncnn/modelbin.h. load_model()provides overloads for file paths,FILE*pointers, raw memory buffers, and customDataReaderobjects, all centralized insrc/net.cpp.- The
const unsigned char*overload enables zero-copy weight loading viaModelBinFromMatArray, directly referencing external memory without duplication. - The internal loading loop (lines 1600–1700 in
src/net.cpp) processes all layers uniformly regardless of the weight source, ensuring consistent behavior across file, memory, and custom data sources. - Unit tests in
tests/test_squeezenet.cppvalidate the memory pointer pathway for embedded deployment scenarios.
Frequently Asked Questions
What is the difference between load_param and load_model in ncnn?
load_param reads the network architecture definition (layer types, dimensions, connections) from a .param file, while load_model reads the actual numeric weights and biases from a .bin file or memory buffer. You must call load_param before load_model so that the network structure is established before the weights are assigned to the layers.
Does ncnn copy model weights when using the external memory overload?
No. When using load_model(const unsigned char*), ncnn constructs a ModelBinFromMatArray that points directly to your buffer. The weight tensors reference your external memory without allocating internal copies, provided the buffer remains valid throughout the lifetime of the ncnn::Net instance.
Can I load ncnn models from encrypted or compressed files?
Yes. By subclassing ncnn::DataReader and implementing the read() and close() methods, you can feed decrypted or decompressed data to load_model(const DataReader&). This allows you to handle decryption logic within the data reader while keeping the network loading code unchanged.
Which load_model method is fastest for mobile deployment?
The load_model(const unsigned char*) overload is typically fastest because it eliminates file system overhead and memory copying. For optimal performance on mobile devices, load the model binary into memory during app startup (or keep it in memory-mapped storage) and pass the pointer directly to ncnn, as demonstrated in the test_squeezenet.cpp benchmark lines 179–201.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →