# How ncnn Manages Model Weights: load_model() vs External Memory Referencing

> Discover how ncnn manages model weights using ModelBin and load_model(). Learn the differences between direct loading and external memory referencing for efficient inference.

- Repository: [Tencent/ncnn](https://github.com/tencent/ncnn)
- Tags: internals
- Published: 2026-02-23

---

**ncnn manages model weights through an abstraction layer called `ModelBin`, allowing the `load_model()` method to ingest weights from disk files, memory buffers, or custom data streams without requiring layers to distinguish between the sources.**

In the Tencent/ncnn inference framework, model parameters—including weights, biases, and quantization tables—are decoupled from the network architecture to support flexible deployment scenarios. The framework provides multiple overloads of `ncnn::Net::load_model` in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp) to handle different data sources, ranging from traditional file I/O to zero-copy external memory referencing. Understanding how ncnn manages model weights through these pathways is essential for optimizing memory usage and loading latency in production environments.

## The ModelBin Abstraction Layer

All layers in ncnn retrieve their weights through the **`ModelBin`** interface, defined in [`include/ncnn/modelbin.h`](https://github.com/Tencent/ncnn/blob/main/include/ncnn/modelbin.h). This abstraction ensures that layer implementations remain agnostic to the underlying storage mechanism.

When a network loads weights, each layer's `load_model` method receives a `const ncnn::ModelBin&` reference. The layer then calls `mb.load(index)` to retrieve tensors:

```cpp
int Convolution::load_model(const ncnn::ModelBin& mb)
{
    weight_data = mb.load(w);
    bias_data   = mb.load(b);
    // ...
}

```

The `ModelBin` implementation handles the actual byte reading. **`ModelBinFromMatArray`** wraps a raw pointer to user-supplied memory (such as a `std::vector<float>`), while the default implementation uses a **`DataReader`** for file-based access. This design allows the same layer code to function correctly whether the weights reside on disk or in RAM.

## Net::load_model() Overloads Explained

The `ncnn::Net` class provides four primary overloads of `load_model()` in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp), all converging on the same internal loop between lines 1600–1700 that iterates over layers and invokes `layer->load_model(mb)`.

### File-Based Loading (Standard I/O)

For deployment scenarios where the model binary (`.bin`) resides on the filesystem, ncnn offers two file-based overloads:

- **`load_model(const char* modelpath)`** ([src/net.cpp#L1838](https://github.com/Tencent/ncnn/blob/master/src/net.cpp#L1838)): Opens the specified path as a `DataReaderFile`.
- **`load_model(FILE* fp)`** ([src/net.cpp#L1832](https://github.com/Tencent/ncnn/blob/master/src/net.cpp#L1832)): Wraps an existing `FILE*` stream in a `DataReaderFile`.

These methods are suitable for desktop applications and standard mobile deployments where the model file is packaged as an app asset.

### Memory-Based Loading (External Referencing)

When the weight data is already present in memory—perhaps decompressed from an archive, received over a network, or stored in a memory-mapped file—the **`load_model(const unsigned char* mem)`** overload ([src/net.cpp#L1923](https://github.com/Tencent/ncnn/blob/master/src/net.cpp#L1923)) provides a zero-copy pathway. This method constructs a `ModelBinFromMatArray` instance that points directly to the provided buffer without allocating internal copies of the weight tensors.

### Custom Data Sources

For encrypted or compressed models, developers can subclass **`DataReader`** and pass it to **`load_model(const DataReader& dr)`** ([src/net.cpp#L1599](https://github.com/Tencent/ncnn/blob/master/src/net.cpp#L1599)). This central routine builds the `ModelBin` from the reader and executes the layer loading loop, enabling on-the-fly decryption or streaming from non-standard sources.

## How External Memory Referencing Works

External memory referencing leverages **`ModelBinFromMatArray`**, implemented in [`src/mat.cpp`](https://github.com/Tencent/ncnn/blob/main/src/mat.cpp) and declared in [`include/ncnn/modelbin.h`](https://github.com/Tencent/ncnn/blob/main/include/ncnn/modelbin.h). When you invoke the pointer overload:

```cpp
const unsigned char* model_data = /* pre-loaded blob */;
ncnn::Net net;
net.load_model(model_data);

```

The framework treats `model_data` as a read-only array of floats and constructs weight matrices that reference this memory directly. This approach eliminates the memory overhead of duplicating large weight files and avoids file system latency during initialization.

The unit tests in [`tests/test_squeezenet.cpp`](https://github.com/Tencent/ncnn/blob/main/tests/test_squeezenet.cpp) demonstrate this pattern in lines 179–201, where the test suite exercises `load_model((const unsigned char*)model_data)` to validate inference without touching the disk.

## Performance Considerations and Use Cases

Choosing the appropriate `load_model` overload depends on your deployment constraints and performance requirements:

- **Standard file deployment**: Use `load_model(const char*)` for simplicity when the `.bin` file is stored on disk or in standard Android/iOS assets.
- **Embedded systems with custom loaders**: Use `load_model(const unsigned char*)` when the model is embedded as a byte array in the executable or loaded into RAM by a custom resource manager.
- **Security-sensitive applications**: Implement a custom `DataReader` subclass for `load_model(const DataReader&)` when models are encrypted or compressed, allowing decryption to occur stream-wise during layer initialization.
- **Low-latency inference**: The external memory pointer overload minimizes initialization time and peak memory usage by avoiding file I/O and buffer copies, making it ideal for real-time applications on resource-constrained devices.

## Practical Implementation Examples

### Loading Weights from Disk

The most common pattern loads network parameters first, then the binary weights:

```cpp
#include "net.h"

ncnn::Net net;
net.load_param("mobilenet_v2.param");
net.load_model("mobilenet_v2.bin");

```

*Source: `Net::load_model(const char*)` ([src/net.cpp#L1838](https://github.com/Tencent/ncnn/blob/master/src/net.cpp#L1838)).*

### Loading from External Memory

For zero-copy loading when the binary data is already in a `std::vector` or mmap buffer:

```cpp
#include "net.h"
#include <vector>

std::vector<unsigned char> buffer = read_from_custom_source();
ncnn::Net net;
net.load_param("shufflenet_v2.param");
net.load_model(buffer.data());  // No copy performed

```

*Source: `Net::load_model(const unsigned char*)` ([src/net.cpp#L1923](https://github.com/Tencent/ncnn/blob/master/src/net.cpp#L1923)).*

### Custom DataReader for Encrypted Models

To load weights from an encrypted blob without writing a temporary file:

```cpp
class EncryptedReader : public ncnn::DataReader {
public:
    EncryptedReader(const unsigned char* key, const unsigned char* data, size_t len)
        : key_(key), data_(data), pos_(0), len_(len) {}

    virtual int read(void* buf, size_t size) {
        decrypt_block(data_ + pos_, buf, size, key_);
        pos_ += size;
        return 0;
    }
    virtual void close() {}

private:
    const unsigned char *key_, *data_;
    size_t pos_, len_;
};

// Usage
EncryptedReader reader(key, encrypted_blob, blob_size);
net.load_model(reader);

```

*Source: `Net::load_model(const DataReader&)` ([src/net.cpp#L1599](https://github.com/Tencent/ncnn/blob/master/src/net.cpp#L1599)).*

## Summary

- ncnn uses the **`ModelBin`** abstraction to decouple layer weight consumption from storage implementations, located in [`include/ncnn/modelbin.h`](https://github.com/Tencent/ncnn/blob/main/include/ncnn/modelbin.h).
- **`load_model()`** provides overloads for file paths, `FILE*` pointers, raw memory buffers, and custom `DataReader` objects, all centralized in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp).
- The **`const unsigned char*`** overload enables **zero-copy** weight loading via `ModelBinFromMatArray`, directly referencing external memory without duplication.
- The internal loading loop (lines 1600–1700 in [`src/net.cpp`](https://github.com/Tencent/ncnn/blob/main/src/net.cpp)) processes all layers uniformly regardless of the weight source, ensuring consistent behavior across file, memory, and custom data sources.
- Unit tests in [`tests/test_squeezenet.cpp`](https://github.com/Tencent/ncnn/blob/main/tests/test_squeezenet.cpp) validate the memory pointer pathway for embedded deployment scenarios.

## Frequently Asked Questions

### What is the difference between `load_param` and `load_model` in ncnn?

`load_param` reads the network architecture definition (layer types, dimensions, connections) from a `.param` file, while `load_model` reads the actual numeric weights and biases from a `.bin` file or memory buffer. You must call `load_param` before `load_model` so that the network structure is established before the weights are assigned to the layers.

### Does ncnn copy model weights when using the external memory overload?

No. When using `load_model(const unsigned char*)`, ncnn constructs a `ModelBinFromMatArray` that points directly to your buffer. The weight tensors reference your external memory without allocating internal copies, provided the buffer remains valid throughout the lifetime of the `ncnn::Net` instance.

### Can I load ncnn models from encrypted or compressed files?

Yes. By subclassing `ncnn::DataReader` and implementing the `read()` and `close()` methods, you can feed decrypted or decompressed data to `load_model(const DataReader&)`. This allows you to handle decryption logic within the data reader while keeping the network loading code unchanged.

### Which `load_model` method is fastest for mobile deployment?

The **`load_model(const unsigned char*)`** overload is typically fastest because it eliminates file system overhead and memory copying. For optimal performance on mobile devices, load the model binary into memory during app startup (or keep it in memory-mapped storage) and pass the pointer directly to ncnn, as demonstrated in the [`test_squeezenet.cpp`](https://github.com/Tencent/ncnn/blob/main/test_squeezenet.cpp) benchmark lines 179–201.