# How CubeVS eBPF Virtual Switch Provides Kernel-Level Network Isolation

> Discover how CubeVS eBPF virtual switch provides kernel-level network isolation. Learn how eBPF and BPF maps enforce policies, filter DNS, and manage SNAT rules in kernel space.

- Repository: [Tencent Cloud/CubeSandbox](https://github.com/TencentCloud/CubeSandbox)
- Tags: how-to-guide
- Published: 2026-07-09

---

**CubeVS achieves kernel-level network isolation by attaching eBPF programs to the Linux Traffic Control (TC) subsystem at three strategic hook points, using high-performance BPF maps to enforce per-sandbox CIDR policies, DNS filtering, and SNAT rules entirely in kernel space before packets reach containerized workloads.**

CubeSandbox leverages CubeVS as its high-performance virtual switch layer to isolate multi-tenant sandboxed workloads. Unlike traditional bridge-based networking that relies on userspace packet processing, CubeVS implements a pure eBPF data plane that intercepts traffic at the kernel's TC layer, eliminating context-switch overhead while enforcing strict network policies that privileged containers cannot bypass.

## eBPF Hook Architecture and Attachment Points

CubeVS installs eBPF programs at three distinct TC hooks to segment traffic flows between the host node, the gateway interface, and individual sandbox TAP devices.

### The Three TC Hook Points

| Hook Point | Source File | Traffic Direction |
|------------|-------------|-------------------|
| `nodenic` | [`nodenic.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/nodenic.bpf.c) | Bidirectional traffic between the host node's physical NIC and the sandbox-side virtual NIC (`cubegw0`). |
| `localgw` | [`localgw.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/localgw.bpf.c) | Egress packets leaving the sandbox toward external networks ("world-side" processing). |
| `mvmtap` | [`mvmtap.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/mvmtap.bpf.c) | Ingress packets entering the sandbox from external sources ("in-gateway" processing). |

According to the CubeSandbox source code in [`CubeNet/cubevs/cubevs.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/CubeNet/cubevs/cubevs.go), these programs are generated at build time using **bpf2go** via `//go:generate` directives. After compilation, each program is pinned under `/sys/fs/bpf` and attached to the appropriate TC qdisc using constants defined at lines 84-89, including `tcHandleClsact` and `tcFilterHandle`.

### Build and Loading Process

The Go control plane in [`cubevs.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/cubevs.go) manages the lifecycle of these eBPF objects. The programs are compiled to BPF bytecode during the build process, then loaded and attached to the TC `clsact` qdisc at runtime. This ensures that policy enforcement occurs at the earliest possible point in the kernel networking stack.

## Data Plane Maps and Structures

CubeVS maintains isolation through a set of specialized BPF maps that store per-sandbox metadata, policy rules, and translation tables. These maps reside in kernel memory and provide constant-time lookups for packet processing.

### Metadata and Lookup Tables

- **`ifindex_to_mvmmeta`** – Stores per-TAP-device metadata including the sandbox's inner IP, UUID, and version. The `UpsertTAPDeviceMeta` function in [`tap.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/tap.go) (lines 58-100) populates this map when creating new sandbox interfaces.
- **`mvmip_to_ifindex`** – Provides reverse mapping from sandbox IPs to their corresponding TAP interface indices, enabling fast destination lookups during packet forwarding.

### Network Policy Enforcement

Network isolation relies on Longest Prefix Match (LPM) Trie maps for efficient CIDR-based filtering:

- **`allow_out_v2`** – An LPM-Trie storing `netPolicyValueV2` structures that whitelist egress destinations. Supports L7-aware flags such as `netPolicyFlagL7Required`.
- **`deny_out`** – An LPM-Trie of `uint32` values implementing blacklist semantics for blocked CIDR ranges.
- **`dns_allow`** – A hash-of-maps structure storing hashed DNS names permitted for each sandbox, checked by DNS tail-call programs (`dns_parse_chunk`, `dns_rev_chunk`).

The `initNetPolicy` function in [`netpolicy.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/netpolicy.go) (lines 86-103) creates these maps on-demand for each sandbox interface, while `newInnerLPMMap` (lines 25-33) allocates the inner LPM-Trie structures used for prefix matching.

### NAT and Port Management

- **`remote_port_mapping`** and **`local_port_mapping`** – Implement static port-NAT for external services, mapping outbound connections to specific host ports.
- **`snat_iplist`** – Maintains a configurable pool of source IPs used for SNAT of outbound traffic, preventing IP spoofing across sandboxes. Management logic resides in [`snat.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/snat.go).

## Packet Flow and Isolation Guarantees

CubeVS enforces isolation by processing every packet through its eBPF data plane before delivery to the sandbox network namespace.

### Ingress Processing (Sandbox to Host)

When a sandboxed workload transmits a packet:

1. The `mvmtap` eBPF program attached to the TAP device intercepts the packet.
2. The program queries `ifindex_to_mvmmeta` to identify the source sandbox.
3. It validates the packet against `allow_out_v2` and `deny_out` LPM-Trie maps, performing longest-prefix matching on destination CIDRs.
4. If DNS traffic is detected, the program checks the `dns_allow` hash map for policy compliance.
5. Upon approval, the program rewrites L2 headers to the host MAC and redirects the packet to the world side using `BPFRedirectFlagIngress`.

### Egress Processing (Host to Sandbox)

For inbound traffic destined to a sandbox:

1. The `localgw` program attached to `cubegw0` intercepts outbound packets.
2. It performs SNAT using the `snat_iplist` map to assign a controlled source IP.
3. The program consults `allow_out_v2` or `deny_out` maps for policy verification.
4. It resolves the destination TAP device via `mvmip_to_ifindex` and redirects the packet to the appropriate interface.

### Node-Level Forwarding

The `nodenic` program handles bidirectional forwarding between the host node's physical NIC and `cubegv0`, managing L2 header rewrites and optional DNS filtering for all node-level traffic.

### Isolation Mechanisms

**Per-Sandbox IP Isolation** – Each sandbox receives a unique `MVMInnerIP` mapped to a specific TAP ifindex. The eBPF maps guarantee that packets can only be delivered to the TAP device owning the destination IP, preventing cross-sandbox sniffing or spoofing.

**CIDR-Based Egress Control** – The LPM-Trie implementation in `allow_out_v2` and `deny_out` enables fine-grained egress policies ranging from specific /32 hosts to broad network ranges, with nanosecond-scale lookup latency.

**DNS Filtering** – The `dns_allow` map stores SHA-hashed DNS names that are validated by dedicated tail-call programs, preventing sandboxes from resolving prohibited domains even when using their own DNS resolvers.

**Kernel-Level Enforcement** – Because all policy decisions occur in eBPF before packets enter userspace or bridge devices, sandboxed workloads running with `CAP_NET_ADMIN` or even root privileges cannot bypass or modify network policies.

## Implementation Details and Source Files

The CubeVS implementation spans multiple Go source files in the `CubeNet/cubevs` package:

- **[`cubevs.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/cubevs.go)** – Central package definition containing map specifications, TC constants (`tcHandleClsact`, `tcFilterHandle`), and the `Params` structure for sandbox configuration.
- **[`tap.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/tap.go)** – TAP device lifecycle management including `UpsertTAPDeviceMeta`, `ListTAPDevices`, and interface lookup functions.
- **[`netpolicy.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/netpolicy.go)** – LPM-Trie creation via `newInnerLPMMap`, policy initialization via `initNetPolicy`, and CIDR parsing utilities.
- **[`snat.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/snat.go)** – SNAT pool management and IP allocation logic.
- **[`miscs.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/miscs.go)** – Helper functions for loading and pinning eBPF objects, constant rewriting, and tail-call orchestration.
- **[`src/localgw.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/src/localgw.bpf.c)**, **[`src/mvmtap.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/src/mvmtap.bpf.c)**, **[`src/nodenic.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/src/nodenic.bpf.c)** – The C source for the eBPF programs compiled to bytecode.

## Practical Usage Examples

The following examples demonstrate common operations using the CubeVS Go API:

```go
// Create a new sandbox TAP device and configure egress policies
params := cubevs.Params{
    MVMInnerIP: net.ParseIP("10.1.2.3"),
    MVMMacAddr: net.HardwareAddr{0x02, 0x42, 0xac, 0x11, 0x00, 0x02},
    // … other required fields …
}
vs, _ := cubevs.New(params)

opts := cubevs.MVMOptions{
    AllowOut:            &[]string{"0.0.0.0/0"},
    L7AllowOut:          &[]string{"10.0.0.0/8"},
    DenyOut:             &[]string{"192.168.0.0/16"},
}
_ = vs.AddTAPDevice(5, net.ParseIP("10.1.2.3"), "sandbox-01", 1, opts)

```

```go
// List all TAP devices currently managed by CubeVS
taps, _ := cubevs.ListTAPDevices()
for _, t := range taps {
    fmt.Printf("TAP %s – IP %s – ifindex %d\n", t.ID, t.IP, t.Ifindex)
}

```

```go
// Remove a sandbox TAP device and clean up associated policy maps
_ = cubevs.DelTAPDevice(5, net.ParseIP("10.1.2.3"))

```

## Summary

- CubeVS provides **kernel-level network isolation** by attaching eBPF programs to TC hooks (`nodenic`, `localgw`, `mvmtap`) rather than using traditional bridge or veth pairs.
- **LPM-Trie maps** (`allow_out_v2`, `deny_out`) enable efficient CIDR-based egress filtering with whitelist and blacklist semantics.
- **Hash-based maps** store per-sandbox metadata (`ifindex_to_mvmmeta`), IP-to-interface mappings (`mvmip_to_ifindex`), and DNS allowlists (`dns_allow`).
- All enforcement occurs in **kernel space** via the eBPF data plane, making policies tamper-proof even for privileged containers.
- The Go control plane in [`cubevs.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/cubevs.go), [`tap.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/tap.go), and [`netpolicy.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/netpolicy.go) manages eBPF object lifecycle, map updates, and policy configuration through a clean API.

## Frequently Asked Questions

### How does CubeVS differ from traditional Linux bridge networking?

Traditional Linux bridges process packets in the kernel's networking stack but rely on userspace daemons or iptables for policy enforcement, introducing context switches and bypass vulnerabilities. CubeVS performs all filtering and forwarding in eBPF programs attached to the TC layer, delivering sub-microsecond latency while ensuring privileged containers cannot circumvent policies because the enforcement occurs before packets enter the sandbox's network namespace.

### What map types does CubeVS use for policy enforcement?

CubeVS uses **LPM-Trie** (Longest Prefix Match) maps for CIDR-based network policies (`allow_out_v2`, `deny_out`), enabling efficient matching of IP addresses against network ranges. It uses **hash maps** for exact-match lookups such as DNS names (`dns_allow`), IP-to-interface translations (`mvmip_to_ifindex`), and TAP device metadata (`ifindex_to_mvmmeta`). The `dns_allow` map implements a hash-of-maps pattern for per-sandbox DNS filtering.

### Can sandboxed containers bypass CubeVS network policies?

No. Because CubeVS attaches eBPF programs to the TC subsystem at the host level, policy enforcement occurs in kernel space before packets reach the sandbox's network namespace or virtual interface. Even containers running with root privileges or `CAP_NET_ADMIN` cannot modify the eBPF programs or maps because they require host-level permissions to access `/sys/fs/bpf` and the TC subsystem, which are controlled by the CubeSandbox runtime.

### How are eBPF programs loaded and pinned in CubeVS?

CubeVS uses **bpf2go** to compile C source files ([`nodenic.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/nodenic.bpf.c), [`localgw.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/localgw.bpf.c), [`mvmtap.bpf.c`](https://github.com/TencentCloud/CubeSandbox/blob/main/mvmtap.bpf.c)) into BPF bytecode at build time. At runtime, the Go control plane loads these objects using the cilium/ebpf library, pins them to `/sys/fs/bpf` for persistence, and attaches them to TC qdiscs using the `clsact` discipline. The constants `tcHandleClsact` and `tcFilterHandle` defined in [`cubevs.go`](https://github.com/TencentCloud/CubeSandbox/blob/main/cubevs.go) (lines 84-89) specify the attachment points for the filters.