How to Implement and Register a Custom ncnn Layer Using `register_custom_layer()`

You can extend Tencent/ncnn with custom operators by subclassing ncnn::Layer, defining creator/destroyer macros, and calling register_custom_layer() before loading any model.

The Tencent/ncnn framework provides an extensible layer registration API that lets you add custom operators without recompiling the core library. The registration interface lives in src/net.h and src/net.cpp, exposing two overloads that accept either a layer type name or a type index. This guide walks through the implementation workflow, references the exact source locations, and highlights the critical pitfalls that cause runtime failures.

Architecture of the Layer Registration System

Understanding the internal structure prevents registration errors and performance bottlenecks.

The Layer Base Class and Registry

All operators inherit from ncnn::Layer, defined in src/layer.h. This base class defines the virtual interface methods load_param(), load_model(), and four possible forward* overloads. The custom-layer registry itself consists of two vectors stored in NetPrivate (internal to src/net.cpp) that hold creator and destroyer function pointers. These are keyed either by a string type name or by an integer type index.

Registration API Entry Points

The Net class exposes two overloads for register_custom_layer():

  • String overload (src/net.cpp, lines 29-74): Accepts a layer type name such as "MyLayer". If the name matches a built-in layer, this call overwrites the built-in implementation; otherwise, it appends a new entry to the custom registry.
  • Integer overload (src/net.cpp, lines 74-118): Works with a raw type index. The implementation masks the custom bit (LayerType::CustomBit) using index & ~LayerType::CustomBit. Supplying an index that corresponds to a built-in layer triggers an overwrite warning.

For C projects, the equivalent functions are ncnn_net_register_custom_layer_by_type() and ncnn_net_register_custom_layer_by_typeindex() in src/c_api.cpp (lines 1470-1494). The Python API provides a thin wrapper around these mechanisms but pre-allocates a fixed capacity of 256 factories (g_layer_factroys in python/src/main.cpp).

Step-by-Step Implementation Guide

Follow this sequence exactly to ensure your custom layer loads correctly.

1. Subclass ncnn::Layer

Create a header and implementation file for your operator. In the constructor, set the boolean flags one_blob_only and support_inplace to declare memory behavior. These flags determine which forward* overload the framework invokes.

// myscale.h
#include "layer.h"

namespace ncnn {

class MyScale : public Layer
{
public:
    MyScale();

    virtual int load_param(const ParamDict& pd) override;
    virtual int load_model(const ModelBin& mb) override;
    virtual int forward(const Mat& bottom_blob, Mat& top_blob,
                       const Option& opt) const override;

private:
    float scale = 1.f;
};

} // namespace ncnn

2. Implement Required Virtual Methods

Implement load_param() to read values from the .param file, and implement the appropriate forward* method based on your flags. If one_blob_only is true, implement forward(const Mat&, Mat&, const Option&).

// myscale.cpp
#include "myscale.h"

DEFINE_LAYER_CREATOR(MyScale)
DEFINE_LAYER_DESTROYER(MyScale)

namespace ncnn {

MyScale::MyScale()
{
    one_blob_only = true;
    support_inplace = true;
}

int MyScale::load_param(const ParamDict& pd)
{
    // Parameter id 0 is the scale factor
    scale = pd.get(0, 1.f);
    return 0;
}

int MyScale::forward(const Mat& bottom_blob, Mat& top_blob,
                     const Option& /*opt*/) const
{
    top_blob = bottom_blob.clone();
    float* outptr = top_blob;
    const int size = top_blob.total();
    for (int i = 0; i < size; ++i)
        outptr[i] *= scale;
    return 0;
}

} // namespace ncnn

3. Register the Custom Layer

Registration must occur immediately after constructing the Net object and before calling load_param() or load_model(). This timing is critical because ncnn resolves layer type names to factory functions during the load phase.

#include "myscale.h"
#include "net.h"

int main()
{
    ncnn::Net net;

    // Method A: Registration by string name
    net.register_custom_layer("MyScale", 
                              MyScale_layer_creator, 
                              MyScale_layer_destroyer);

    // Method B: Registration by type index (faster, avoids string lookup)
    // net.register_custom_layer(ncnn::layer_to_index("MyScale"),
    //                           MyScale_layer_creator,
    //                           MyScale_layer_destroyer);

    net.load_param("model.param");
    net.load_model("model.bin");
    
    // Inference code follows...
}

The DEFINE_LAYER_CREATOR and DEFINE_LAYER_DESTROYER macros generate function pointers with the exact signatures required: ncnn::Layer* (*creator)() and void (*destroyer)(ncnn::Layer*).

Common Pitfalls When Using register_custom_layer()

Avoid these errors that frequently appear in production code and issue trackers.

Accidentally Overwriting Built-in Layers

The string-based overload looks up the type name via layer_to_index(). If you pass a name like "Convolution" or "ReLU", the call silently overwrites the built-in implementation. Verify that your custom layer uses a unique type name, or intentionally overwrite only after understanding the consequences.

Registering After Model Loading

If you call register_custom_layer() after load_param(), ncnn throws an unknown layer type error because the layer graph has already been constructed. The framework resolves type names to factory pointers during the load phase, not during inference.

Mismatched Creator/Destroyer Signatures

The macros expect specific C++ linkage. If you manually write creator functions instead of using DEFINE_LAYER_CREATOR, ensure the signature exactly matches ncnn::Layer* creator(). Similarly, custom destroyers must accept ncnn::Layer* and return void. Signature mismatches produce cryptic template instantiation errors.

Incorrect Forward Overload Selection

Setting one_blob_only = true but implementing forward_inplace() or a multi-blob forward causes the framework to trigger fallback paths or crash. Set these flags in the constructor to match the virtual method you actually implement, as detailed in the step-by-step guide at docs/developer-guide/how-to-implement-custom-layer-step-by-step.md.

Custom Index Out of Range

When using the integer overload, manually supplying an index that lacks the custom bit mask can corrupt the registry or trigger built-in overwrites. Always use ncnn::layer_to_index("YourLayer") or let the string overload assign the index automatically.

Python Binding Capacity Limits

The Python API stores factories in a fixed-size array (g_layer_factroys). Registering more than the default 256 custom layers raises a runtime error. Recompile the Python bindings with a larger NCNN_MAX_CUSTOM_LAYER value if necessary.

Thread Safety Violations

Registration modifies global vectors inside NetPrivate. Calling register_custom_layer() from multiple threads simultaneously creates race conditions. Perform all registrations in a single thread before spawning inference workers.

Missing Destroyers for GPU Resources

While optional for CPU-only layers, custom layers that allocate Vulkan objects (VkBuffer, VkImage) must provide a destroyer via DEFINE_LAYER_DESTROYER to release GPU memory before the Net destructor runs. The default delete does not handle device-specific cleanup.

NCNN_STRING Guard Issues

If ncnn is compiled with NCNN_STRING=0, the string-based registration API is disabled. Code using the string overload will fail to compile. Use the integer overload (register_custom_layer(index, creator, destroyer)) for embedded builds that disable string support.

Summary

  • Inherit from ncnn::Layer and implement load_param, load_model, and the correct forward* method based on one_blob_only and support_inplace flags.
  • Use macros: DEFINE_LAYER_CREATOR and DEFINE_LAYER_DESTROYER generate the function pointers required for registration.
  • Register early: Call net.register_custom_layer() immediately after Net construction and before any load_param() or load_model() calls.
  • Avoid collisions: Use unique layer type names to prevent overwriting built-in operators.
  • Respect limits: Remember the 256-layer capacity in Python bindings and the thread-safety requirements of the registry.

Frequently Asked Questions

Can I override a built-in ncnn layer with my own implementation?

Yes. If you pass a built-in type name (e.g., "Convolution") to the string overload of register_custom_layer(), the function overwrites the existing factory pointer in the registry. This behavior is defined in src/net.cpp lines 29-74. Use this capability carefully, as it affects all subsequent inference runs using that Net instance.

Why does my model fail to load with "unknown layer type" even after registration?

This error occurs when load_param() executes before register_custom_layer(). The framework resolves layer type names to integer indices during the parameter loading phase. You must register all custom layers immediately after constructing the ncnn::Net object, as demonstrated in examples/yolox.cpp and tests/test_squeezenet.cpp.

Do I need to provide a custom destroyer if my layer only uses CPU memory?

No. If your layer does not allocate GPU resources or other non-trivial assets, you can pass nullptr for the destroyer argument or use the default DEFINE_LAYER_DESTROYER macro, which maps to standard delete. However, for Vulkan-accelerated layers, you must implement a custom destroyer to release Vk* objects before the network cleanup completes.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →