How the Internal Representation of IP Ranges Affects Lookup Accuracy and Performance in geoip

The loyalsoldier/geoip tool stores IP ranges inside a netipx.IPSet structure, providing exact CIDR matching through radix-tree-based O(log n) containment checks that deliver sub-microsecond query latency even with millions of prefixes.

The internal representation of IP ranges directly determines how accurately and quickly a geo-IP database can answer containment queries. In the open-source loyalsoldier/geoip repository, every CIDR block is normalized and stored within highly optimized immutable sets that guarantee precise boundary matching while maintaining logarithmic-time lookup performance.

Building Canonical IP Sets from Raw Input

The transformation from raw CIDR strings to query-ready structures happens inside lib/entry.go, where strict normalization ensures no ambiguity enters the dataset.

Parsing and Normalization via processPrefix

The processPrefix function (lines 70–138) accepts diverse input types—including net.IP, net.IPNet, netip.Prefix, and strings—and converts them into canonical netip.Prefix values. This function actively rejects IPv4-mapped IPv6 addresses (e.g., ::ffff:192.0.2.1) by returning ErrInvalidCIDR, preventing address family contamination. All valid addresses undergo unmapping via addr.Unmap(), ensuring IPv4 addresses remain 32-bit values rather than being stored as 128-bit IPv6 equivalents.

Separate Builders for Each Address Family

Inside the add function (lines 49–64), the code detects the IPType (IPv4 or IPv6) and allocates a dedicated netipx.IPSetBuilder for each family. This separation prevents cross-family pollution and allows the final lookup logic to query only the relevant dataset. When entries are merged in lib/container.go, the AddSet method deduplicates overlapping prefixes and coalesces adjacent ranges, producing the minimal node count necessary for optimal traversal.

Radix-Tree Structure Enables O(log n) Lookups

The immutable IPSet generated by buildIPSet (lines 307–324) utilizes a radix-tree-like structure internally. This representation transforms the lookup operation from a linear scan into a logarithmic search.

Immutable Set Construction

Once builder.IPSet() is invoked, the resulting structure becomes read-only and highly optimized for cache-friendly traversal. The builder is discarded after construction, eliminating further allocations or re-sorting during the query phase. This immutability guarantees constant-time memory layout and prevents race conditions during concurrent lookups.

Containment Testing in lib/container.go

The lookup chain flows from lookupCmd (CLI) through special.Lookup to container.Lookup, which ultimately calls ipset.Contains(addr) or ipset.ContainsPrefix(prefix) inside lib/container.go. Both methods execute in O(log n) time relative to the number of stored prefixes, enabling the tool to process millions of CIDR blocks while maintaining sub-5-microsecond latency on modest hardware.

Accuracy Guarantees Through Strict Validation

The internal representation enforces exact matching semantics that prevent false positives common in less rigorous implementations.

Exact CIDR Boundary Preservation

Unlike simple range comparisons that might treat 192.0.2.0/24 as matching 192.0.2.0/23, the netipx.IPSet retains the exact mask length specified during input. The ContainsPrefix method verifies that the query prefix falls entirely within a stored CIDR block, ensuring strict hierarchical matching without unintended parent-block collisions.

Unified Address Normalization

By rejecting IPv4-mapped IPv6 strings and unmapping all addresses before storage, processPrefix eliminates a class of subtle bugs where IPv4 addresses might be incorrectly matched against IPv6 ranges or vice versa. This normalization guarantees that the radix tree contains canonical representations only, making the lookup results deterministic across different input formats.

Programmatic Example: Building and Querying Sets

Below is a complete example demonstrating how the internal representation handles range construction and querying:

package main

import (
    "fmt"
    "github.com/Loyalsoldier/geoip/lib"
    "github.com/Loyalsoldier/geoip/plugin/plaintext"
)

func main() {
    // Create a fresh container
    c := lib.NewContainer()

    // Build an entry called "CN" from CIDR lists
    entry := lib.NewEntry("CN")
    _ = entry.AddPrefix("1.0.0.0/24")   // IPv4
    _ = entry.AddPrefix("2001:db8::/32") // IPv6

    // Add to container
    _ = c.Add(entry)

    // Lookup an IPv4 address
    if lists, ok, _ := c.Lookup("1.0.0.5"); ok {
        fmt.Println("Found in:", lists) // → Found in: [CN]
    }

    // Lookup an IPv6 address
    if lists, ok, _ := c.Lookup("2001:db8::1"); ok {
        fmt.Println("Found in:", lists) // → Found in: [CN]
    }
}

All radix-tree construction, prefix merging, and fast containment testing happens transparently within the netipx library dependency declared in go.mod.

Summary

  • Storage Structure: IP ranges live inside netipx.IPSet objects built via IPSetBuilder, with separate instances for IPv4 and IPv6 to prevent cross-family contamination.
  • Performance Characteristics: The radix-tree implementation provides O(log n) containment checks (Contains, ContainsPrefix) that maintain sub-microsecond latency across millions of prefixes.
  • Accuracy Mechanisms: Strict normalization in processPrefix rejects ambiguous formats, unmaps addresses to canonical forms, and preserves exact CIDR boundaries to prevent false positives.
  • Immutability Benefits: Once built via buildIPSet, the read-only structure offers cache-friendly traversal and eliminates runtime allocations during queries.

Frequently Asked Questions

Why does geoip use separate builders for IPv4 and IPv6?

Separating address families into distinct IPSetBuilder instances prevents IPv4-mapped IPv6 addresses from contaminating IPv4 datasets and allows the lookup logic to query only the relevant tree. This separation happens in lib/entry.go during the add phase and ensures that 192.0.2.1 never matches against an IPv6 range containing ::ffff:192.0.2.1.

How does the radix tree handle overlapping CIDR blocks?

When entries are merged in lib/container.go, the AddSet method automatically deduplicates overlapping prefixes and coalesces adjacent ranges into the minimal set of nodes required to represent the coverage. This normalization reduces memory footprint and guarantees that the tree depth—and therefore the lookup time—remains logarithmic relative to the unique prefix count.

What prevents IPv4-mapped IPv6 addresses from causing lookup errors?

The processPrefix function in lib/entry.go explicitly detects and rejects strings like ::ffff:192.0.2.1 by returning ErrInvalidCIDR. Additionally, all valid addresses undergo Unmap() conversion, ensuring IPv4 values are always stored as 32-bit addresses rather than 128-bit IPv6 equivalents, which eliminates cross-family matching bugs.

Is the IP range lookup performance consistent regardless of database size?

Yes. Because netipx.IPSet uses a radix-tree-like structure, both Contains and ContainsPrefix operations execute in O(log n) time relative to the number of stored prefixes. Even with millions of CIDR blocks, the tool maintains sub-5-microsecond query latency because the tree depth grows logarithmically rather than linearly with the dataset size.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →